cron · Aug 18, 2026
The Nightly Export That Ran Twice
A cron job that took 70 minutes on an hourly schedule produced duplicate rows for three days before anyone noticed. The timeline, the fix, and a lock you can run yourself.
This was a small incident with a boring cause, which is the best kind to write down. Nobody paged anyone. Nothing went down. A report was wrong for three days.
What happened
The export job copies the previous day's orders into a reporting table. It's still called the nightly export, a name from before finance asked for fresher numbers and we moved it to an hourly schedule. At that point it finished in about four minutes. Over eighteen months the orders table grew, the job slowed, and one Tuesday it crossed an hour.
From then on, whenever a run overlapped the next, two copies of the job read the same unexported rows and both inserted them. The reporting table had no uniqueness constraint, so the database accepted everything without complaint.
Timeline, all times local:
Tue 01:00: run A starts, as it has for a year and a half
Tue 02:00: run A is still going, run B starts on top of it
Tue 02:10: run A finishes at 70 minutes, run B has already read half the same rows
Wed 09:20: finance asks why revenue for Tuesday is higher than the payment processor says
Thu 10:05: we find the overlap by lining up job start times in the scheduler log
Thu 11:40: lock deployed, duplicates deleted, report rerun
The Thursday gap is the part I'd like to shrink. Two people had looked at the same wrong number on Wednesday and each assumed the other had checked the job.
The fix
We did two things, and the order matters. First, stop the overlap. The lock is a directory, because mkdir either creates it or fails, and it can't do both at once:
#!/bin/bash
# Skip this run if the previous one is still going. mkdir is atomic.
LOCK=/tmp/nightly-export.lock
if ! mkdir "$LOCK" 2>/dev/null; then
echo "$(date +%T) previous run still active, skipping"
exit 0
fi
trap 'rmdir "$LOCK"' EXIT
echo "$(date +%T) export started"
sleep 3
echo "$(date +%T) export finished"I started it, waited a second, and started it again. The second copy printed the skip line and exited cleanly while the first carried on:
04:11:57 export started
04:11:58 previous run still active, skipping
04:12:00 export finishedThe trap is what releases the lock, including when the job fails. A hard kill skips traps, so a stale lock can survive a crashed machine. For a job that can wait for a human to look, that's an acceptable failure. For one that can't, store a PID in the lock and check it.
Second, make the data refuse duplicates: a unique constraint on the order id in the reporting table. The lock stops this particular cause, and the constraint stops the next cause nobody has thought of yet.
What we changed besides code
The scheduler now logs run duration, and an alert fires when any job exceeds half of its interval. Ours would have warned us eight months before the first overlap.
We also added a line to the finance checklist: when a number looks wrong, write in the channel who is investigating. Whoever opens the thread owns it until they close it.
No comments yet
Comments are open. Have a thought or a question? Share it below.