We run unattended coding agents overnight in 45-minute blocks. They take the top item off a queue, do it, verify it, commit, and stop. In the morning you read what they did.
One morning the repository had nine new commits of genuinely good work, and was missing four of ours.
Not conflicted. Missing. The remote's history did not contain them, and nothing had complained.
How it looks from inside
The agent had branched from a base older than the current master. It worked, tested, committed and pushed — all correctly, all on top of a tree that predated four commits made the previous evening. One of those was a fix for a ceiling constant that had made six of seven game levels uncrossable. That bug was now back in master.
Here is the part that makes this worse than a merge conflict. The agent's own quality gates passed, and they were right to. It ran the difficulty harness and got 40 wins against a recorded baseline of 40. Green. But the baseline had been raised to 85 in one of the missing commits, and the constant it was measuring had been fixed in another. It was measuring the old game against the old number and getting the old, correct answer.
Every signal was honest. Every signal was useless.
Why nothing caught it
A merge conflict is loud, and that is its best feature. It stops you. A stale base produces no conflict at all when the two sets of changes touch different files, or when the newer work is simply absent from the branch you are pushing. Git has nothing to object to. You asked to make the remote look like your branch, and it did.
Our safeguards were aimed at the wrong failure. We had a lock so two agents could not edit one worktree at once, which is a real problem and it worked. We had a rule against force-pushing, which held. We did not have anything asserting that work already shipped was still present.
The check
One command, and it is cheap enough to run constantly:
git merge-base --is-ancestor <commit> origin/master
Exit zero means the commit is in the remote's history. Non-zero means it is not. Feed it the commits you shipped recently and you find out immediately, instead of when someone reports the old bug.
Two habits around it. Agents should branch from a freshly fetched master at the start of every block, not from whatever the worktree happened to be on. And the recovery is a merge, never a reset — both sides had real work, and the merge turned out to be three trivial conflicts: two cache-busting hashes and a build identifier.
The wider version of this
We found three dead pipelines in one day, and they had the same shape. A deploy step that stopped firing, so nine commits sat unshipped for seven hours. A blog that went nine days without a post. A social queue with 82 items pending and 42 expired unsent, while its timer succeeded every fifteen minutes. All green on our own dashboard, which reported liveness: next run, last run, exit code.
Most scheduled work fails by producing nothing rather than by crashing. A queue drainer that drains nothing exits zero. A writer that writes nothing exits zero. An agent that branches wrong passes its tests.
So the panel now asks a different question. Not is the job alive but did the artefact move: a commit added, a post actually sent, a log that grew, the live build identifier matching the repository. Each check carries a staleness budget and the reason for that number, and a check that cannot run reports unknown rather than green. Being green on missing data is the failure, not a nuisance.
The first run of it flagged the blog at 210 hours against a 168-hour budget, and the queue as degraded. Both were true, both had been true for over a week, and both had been invisible.
We build automation and then find out how it lies to us. Rebel Studios.
