Why I Add Small Delays to UI Tests (and the one time it nearly doubled CI time)

I started inserting tiny, targeted delays in end‑to‑end UI tests to expose race conditions caused by real‑world latency. It found bugs—until it slowed our CI and masked one timing bug.

Written by: Arjun Malhotra

Person typing on a laptop with code visible on the screen and a coffee cup to the side
Photo by Brooke Cagle on Unsplash

I was on a client call. The demo worked on my laptop. It failed for them. They were on a corporate Wi‑Fi in Pune with a laptop that felt like it was from 2016. Their session kept missing a UI update and clicking an inactive button. The logs showed nothing. Locally, headless Playwright ran the same flow in 400ms and passed. The customer saw 1,200–1,800ms delays and a race.

That was my friction: tests that were too “fast” to find bugs that users on flaky networks actually hit. We had an occasional support ticket every few weeks. Each one was a lottery: reproduce on a physical phone, boot a test VM, pray. I needed a reliable way to make our automated tests behave more like real users in India — not just fast CI bots on an infinite pipe.

What I changed (and why it matters) I started adding deliberate, tiny delays in specific spots of our E2E tests: between sending an action and asserting the resulting UI state. Not big sleeps. Controlled, short delays — often 100–200ms, sometimes a 500ms jitter in network‑heavy paths like payments or SSO. The goal was not to slow everything, but to let the browser and the app’s background work (renders, async state reconciliation, mobile network jitter) catch up in the same way a real device would.

Concretely:

Why this worked Two things were happening in production that our original tests never simulated:

  1. Browsers on cheap Indian devices and corporate proxies reorder rendering / repaint under load. A 40ms paint difference can change whether a button becomes clickable in time.
  2. Payment PSPs and banks in India often respond with variable latency. Our fast tests never hit the slower tails, so race conditions in our front‑end state machine stayed hidden.

The delays exposed multiple real bugs in a week:

Failure: it nearly doubled our nightly CI This change found bugs so effectively we turned the delays on for our nightly runs. And then I made a dumb decision: I enabled the jitter globally across the suite to “be safe”. Nightly CI time jumped from 90 minutes to 160 minutes. That’s not a headline metric — 70 extra minutes every night is developer friction, delayed feedback, and cost (our hosted runners bill climbs).

Worse, one of the bugs we found was masked by the delays. A fragile timing bug in an optimistic update only reproduced when the round‑trip was <70ms. Our deliberate delays made the test pass consistently, hiding that specific regression. We had a week of false confidence until an on‑call customer hit it.

Tradeoffs I accepted (and how I recovered)

A few practical rules I settled on

India specifics that matter

What I walked away with Tiny, deliberate waits are not a band‑aid. They’re a diagnostic tool: they make hidden races visible and reproducible. But they’re also a blunt instrument — used everywhere they become a cover-up and a CI tax. The rule I actually kept: make realism targeted, toggleable, and measured. That single policy turned intermittent support calls into reproducible bugs I could fix in a morning, without permanently doubling our CI bill.