How to Manage Cache and Optimize Build Performance

Intermediate 20 minutes DevOps / Platform Engineers Advanced Configuration

Monitor cache hit ratios, understand the blob storage architecture, purge cached packages, and optimize cache performance for faster builds.

Overview

Chainsaw’s pull-through cache is one of its most impactful features for developer productivity. By caching packages locally, builds avoid redundant upstream fetches, reduce latency, and continue working even when upstream registries experience outages. This tutorial covers monitoring, managing, and optimizing the cache.

Prerequisites

  • Admin or Owner role in Chainsaw
  • Traffic flowing through the proxy (cache builds up over time)

Step 1: Understand Cache Architecture

Chainsaw stores cached artifacts in a sharded blob store on the filesystem:

/data/blobs/
  ├── ab/
  │   ├── cd1234...  (lodash-4.17.21.tgz)
  │   └── ef5678...  (express-4.18.2.tgz)
  ├── bc/
  │   └── ...
  └── ...

Artifacts are content-addressed (stored by hash) to avoid duplicates and enable integrity verification.

Sharded blob store architecture
Cached artifacts are stored in a sharded, content-addressed blob store

How Caching Works

  1. First request: Chainsaw fetches from upstream, stores in blob store, serves to client
  2. Subsequent requests: Served directly from blob store (cache hit)
  3. Metadata: Package metadata is cached separately in the database

Step 2: Monitor Cache Hit Ratio

Navigate to the Overview dashboard. The Cache Hit Ratio KPI shows the percentage of requests served from cache vs fetched from upstream.

Cache hit ratio metric
Monitor your cache hit ratio — higher is better for build performance

Interpreting the Ratio

RatioAssessment
90-100%Excellent — most packages served from cache
70-89%Good — cache is working well, new packages pulling it down
50-69%Fair — consider if teams are installing many unique packages
< 50%Poor — cache may be too small, or many first-time installs
A low cache hit ratio is normal in the first weeks after deployment as the cache warms up. It improves as more packages are requested and cached.

Step 3: View Per-Repository Cache Statistics

Navigate to Repositories and click into a specific repository to see its cache performance:

  • Cached package count
  • Cache size
  • Hit ratio for this repository
  • Most frequently accessed packages
Per-repository cache statistics
Each repository shows its own cache hit ratio and package count

Step 4: Purge Cached Packages

Sometimes you need to clear cached artifacts — for example, when a malicious version needs to be removed from the cache, or when storage is running low.

Purge from the BOM Page

Navigate to Bill of Materials, find the package, and use the Purge Cache control.

Cache purge controls in the BOM
Purge specific packages from the cache via the Bill of Materials page

When to Purge

ScenarioAction
Malware detected in cached packagePurge immediately
Upstream released a corrected versionPurge old version, re-fetch
Storage capacity concernsPurge oldest/least-used packages
Corrupted artifactPurge and re-fetch
Purging removes the cached artifact. The next request for that package will fetch from upstream, temporarily increasing latency for that request.

Step 5: Optimize Cache Performance

Storage Provisioning

Estimate your cache storage needs based on your ecosystem usage:

EcosystemTypical Package Size1000 Packages
npm50KB - 5MB~500MB
PyPI100KB - 50MB~2GB
Maven100KB - 100MB~5GB
Docker50MB - 2GB~100GB+
Go10KB - 10MB~500MB
Docker images are by far the largest. If storage is constrained, consider enabling Docker proxying only for the most frequently used base images.

Network Topology

For optimal performance, deploy Chainsaw close to your build infrastructure:

  • Same datacenter/region as CI/CD runners
  • Low-latency network between Chainsaw and build machines
  • High-bandwidth connection to upstream registries (for cache misses)
Optimal network topology
Deploy Chainsaw close to your build infrastructure for lowest latency

Warm the Cache Proactively

For critical packages, you can warm the cache before builds by running installs:

# Warm npm cache
npm install --prefer-online --registry https://chain305.com/chainproxy/repository/@default/npmjs/

# Warm pip cache
pip download -r requirements.txt \
  --index-url https://chain305.com/chainproxy/repository/@default/pypi/simple/ \
  --dest /tmp/pip-cache

Step 6: Monitor Upstream Availability

When upstream registries go down, Chainsaw serves cached packages transparently. Monitor for upstream failures:

  • Check server logs for upstream connection errors
  • Watch for increased 404s on uncached packages
  • Monitor cache hit ratio — it rises during upstream outages (only cache hits succeed)
Cache serving during upstream outage
During upstream outages, cached packages continue serving — only uncached packages fail

Step 7: Database vs Blob Store

Chainsaw stores data in two layers:

LayerStoresTechnology
DatabaseMetadata, policies, users, events, package infoManaged relational database
Blob StoreActual package artifacts (tarballs, JARs, wheels)Filesystem (sharded)

For scaling:

  • Single instance: database + local filesystem
  • Multi-instance: database + shared filesystem (NFS) or object storage
For Kubernetes deployments, use a shared persistent volume for the blob store and a managed database instance to enable horizontal scaling.

Best Practices

PracticeBenefit
Monitor cache hit ratio weeklyEarly detection of issues
Purge known-malicious packages immediatelyPrevent re-installation from cache
Size storage for your largest ecosystemAvoid cache eviction
Deploy close to build infrastructureMinimize network latency
Use lockfiles in buildsConsistent caching behavior
Warm cache after deploymentFirst builds don’t hit cold cache

Next Steps