A tiny on‑demand S3 proxy that stopped my CI chewing bandwidth — and the stale file that taught me cache humility

I built a small S3 proxy (₹300 VPS) that cached large test fixtures for CI and local dev, cutting builds and egress costs — until a stale object made a test lie.

Written by: Arjun Malhotra

A laptop on a wooden desk with a notebook and a cup of coffee
Photo by Firmbee on Unsplash

It was 1:10 AM and our CI had been running for two hours. The same job was downloading a 500 MB fixture from S3 on every new runner — hitting our hosted CI egress repeatedly, failing sometimes because the office connection dropped, and burning through the team’s tiny monthly cloud budget. My laptop was on the meet-calling everyone to sleep tone in my head.

We had three large test fixtures (media blobs and a synthetic dataset) that developers and CI repeatedly pulled. Each CI worker re-downloaded them for every job because we didn’t have a shared cache, and our self-hosted runners were edge-located with flaky upstream links over a Bengaluru office NAT. I needed a cheap, reliable cache that behaved like S3 for GETs and required near-zero changes to test harnesses.

What I built: a tiny HTTP proxy that answers S3 GETs, caches objects to disk, serves Range requests, respects If-None-Match/ETag, and falls back to the real S3 when it can’t be reached. I deployed it to a ₹300/month VPS (one of those small DigitalOcean/Hetzner machines you see in side‑projects). Pointing CI at a single S3 endpoint (S3_ENDPOINT=http://proxy.internal) was the only config change.

Why it worked

How I deployed it without messing with S3 or tests

The day it lied to us

A week in, a flaky integration test started passing in CI but failing locally for a couple of engineers. Local runs were pulling a freshly generated fixture (with a bug we’d fixed), but CI kept using the old one. The proxy had cached the old file and served it until its TTL expired. The team had been overwriting the same S3 key in place (copying a new object over an old key), which in our workflow didn’t change the ETag reliably because of how the upload was done.

Result: CI was testing the old behaviour. We shipped a PR that relied on the new fixture and it passed CI but failed in a customer scenario.

What I learned (the hard rules)

Tradeoffs and honest constraints

The takeaway I actually walked away with

If your CI or dev workflow re-downloads the same large blobs repeatedly, a small on‑demand caching proxy is a pragmatic win — cheap, fast to implement, and immediately tangible in India where bandwidth and egress costs matter. But treat cached objects as immutable unless you design for mutations: version keys, set clear TTLs, and monitor the cache. That one extra 24‑hour TTL saved me a lot of build time until it nearly shipped a lie.