testing · Aug 26, 2026
Hunting a Test That Fails One Run in Fifty
A flaky test is a failure you can't replay. Pin the randomness behind a seed, sweep the seeds, and the one-in-fifty ghost becomes a number you can type into a terminal.
The test had failed three times in two weeks. Each time somebody re-ran the pipeline, it went green, and each time somebody shrugged. That shrug has a cost. After the third one, nobody trusted a red build from that suite, and a red build you don't trust is how real regressions ship.
Why you can't just "run it again"
A flaky test fails for a reason. The reason depends on something the test doesn't control: timing, ordering, a clock, a random number. Re-running only rolls the dice again. What you need is a way to hold the dice still.
So the first job isn't finding the bug. It's making the failure repeatable. Every minute spent reading code before that point is a guess.
Make the randomness a parameter
Our test shuffled the finishing times of two writers and checked that neither lagged the other too far. The randomness came from Math.random(), which can't be replayed. The fix is to take the random source as an argument and hand it a seeded one. This is a small PRNG, written inline so there's no dependency:
import { test } from "node:test";
import assert from "node:assert/strict";
function rng(seed) {
return () => {
seed = (seed + 0x6d2b79f5) | 0;
let t = Math.imul(seed ^ (seed >>> 15), 1 | seed);
t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}
function lag(rand) {
const a = Math.floor(rand() * 11);
const b = Math.floor(rand() * 11);
return a - b;
}
test("writer A never trails writer B by more than 7 ms", () => {
const seed = Number(process.env.SEED ?? 1);
assert.ok(lag(rng(seed)) <= 7, `SEED=${seed}`);
});The assertion message carries the seed. That one detail matters more than the rest: when CI fails, the log tells you exactly which dice to throw.
Sweep, then replay
With a seed in hand, a loop turns "fails sometimes" into a list:
for s in $(seq 1 100); do
SEED=$s node --test flaky.test.mjs >/dev/null 2>&1 || echo "seed $s fails"
doneOn my machine that printed two lines, seed 44 fails and seed 89 fails. Two in a hundred is close to the real-world rate we'd been seeing, and the sweep took under a minute. Then SEED=44 node --test flaky.test.mjs fails identically every time I run it, which is the whole point. A bug that fails on demand can be stepped through, printed, bisected and fixed.
Here the cause is plain once you can see it: the assertion tolerates a gap of 7 ms, and the writers' timings can disagree by up to 10. Seed 44 simply draws a large first number and a small second one. Either the tolerance is wrong or the writers need an ordering guarantee. That's a design decision, which is a better thing to argue about than a ghost.
Keep the seed in CI
We left the seeded runner in place and made CI choose a fresh seed per run, printing it first. If a failure appears, the log holds the seed, and anyone can paste it into a terminal. The ghost now leaves a name and address.
No comments yet
Comments are open. Have a thought or a question? Share it below.