# Nx Remote Cache Server

Why rebuild what hasn't changed? A custom cache server that slashed CI times and gave developers their coffee breaks back.

Source: https://miloscvetkovic.dev/work/nx-remote-cache

- Tagline: faster CI builds
- Category: DEVOPS
- Status: PRODUCTION
- Metric: 5× faster builds
- Basis: CI build time with the remote cache, against the same pipelines rebuilding the entire monorepo on every run.
- Tags: Bun, Elysia, Azure Blob Storage
- Published: 2026-09-09
- Updated: 2026-10-01

## The Challenge

Every CI run rebuilt the entire monorepo, however small the change that set it off. No run reused what an earlier run had already built, so code nobody had touched got compiled again, and again, and again. Pipelines ran long enough to fit a coffee break, which meant developers waited, and the full test suite took so long that running all of it felt like a luxury rather than a habit. Meanwhile the cloud bills climbed, because every one of those rebuilds ran on compute we paid for. "Works on my machine" was still something people said with a straight face. The math was simple: we were paying to compile the same unchanged code hundreds of times a day. That left one obvious question worth answering properly: why rebuild what hasn't changed?

## My Approach

I built a cache server from scratch, on Bun for raw speed, with Elysia handling the requests, and made it speak Nx's built-in self-hosted remote cache protocol. The design is two-tier caching. The hot tier is an LRU cache in memory, capped by default at 100 entries, 500MB in total and 10MB per artifact: frequently accessed artifacts up to that size stay there, and the least recently used make room when it fills. The cold tier is Azure Blob Storage, the persistent cache, so artifacts outlive any one server process. A cache that CI leans on also has to be hard to abuse, so access runs on two tokens, each checked with a timing-safe comparison so response times give nothing away about a wrong guess. The read token can only download; the write token can also upload, and only pushes to main hold it, so a pull request cannot poison the cache. A rate limit of 1000 requests a minute per IP keeps it stable, and health checks integrate it with the container orchestration on Azure Container Apps. A health probe with a 5-second timeout runs before the Nx steps, and a server that does not answer it leaves the run on its local cache. The rule underneath fits on one line: if it hasn't changed, we don't rebuild it. Period.

## How It Works

1. Before its Nx steps, a CI run probes the cache server with a 5-second timeout; if the server does not answer, the run carries on with its local cache.
2. When the run reaches a part of the monorepo whose code hasn't changed, Nx asks the server for the stored artifact by its hash (GET /v1/cache/:hash) instead of compiling that code again.
3. The server checks the read or write token with a timing-safe comparison and applies a rate limit of 1000 requests a minute per IP.
4. The server looks in its in-memory LRU cache first, which by default holds frequently accessed artifacts of up to 10MB each.
5. If memory doesn't have the artifact, the server falls back to Azure Blob Storage, the persistent cold tier.
6. Only when neither tier has it does the run build that code. A push to main, which holds the write token, then uploads the new artifact (PUT /v1/cache/:hash); pull-request runs hold only the read token and never write.
7. An entry never changes once written: a second upload for the same hash gets 409 Conflict, and Blob Storage deletes artifacts after 14 days.
8. The server itself is built on Bun and Elysia and runs on Azure Container Apps, with a liveness endpoint that answers while the process is up and a readiness endpoint that also checks storage.

## Key Contributions

- Built LRU in-memory caching for frequently accessed artifacts
- Implemented Azure Blob Storage backend for persistent cache
- Added dual-token authentication (read/write) with timing-safe comparison
- Configured rate limiting (1000 req/min) for stability
- Created health checks for container orchestration integration

## Impact

- CI pipelines went from coffee-break length to near-instant
- "Works on my machine" became "works everywhere, identically"
- Cloud compute bills dropped noticeably
- Developers actually run the full test suite now (because it's fast)

## Tech Stack

| Category | Items |
| --- | --- |
| Runtime | Bun |
| Framework | Elysia |
| Storage | Azure Blob Storage, LRU Cache |
| Auth | Dual-token system |
| Infrastructure | Azure Container Apps |
