One morning our site was nine commits behind. The overnight worker had committed and pushed all of them. The push succeeded. The tests passed. The deploy step reported success. Users had seen none of it for seven hours.
Two facts explain it, and both are easy to get wrong.
A push deploys nothing unless something is listening
Our service had no source connected. No repository, no image. You can see it plainly in the config — there is simply no source key:
{"config": {
"build": {"builder": "RAILPACK"},
"deploy": {"runtime": "V2"},
"networking": {"customDomains": {"example.com": {}}}
}}
Every deploy that had ever happened was somebody typing railway up. That works perfectly and it is invisible: the service is live, the domain resolves, health checks pass. Nothing about the running site tells you that pushing to the default branch does not reach it.
When the automation that was calling railway up stopped, there was no second mechanism, and nothing reported the absence. A deploy that never starts produces no failed build to notice. Our deployment list showed the last success and then nothing at all — not one FAILED, not one CRASHED. Just a gap.
It uploads the folder, not the commit
This is the one to internalise before you automate anything. railway up tars up the working directory and sends that. Not HEAD. Not the branch. The directory, as it exists on disk, at that moment.
Most of the time the difference does not bite, because you deploy from a clean tree. It bites when a human is mid-experiment in that folder, or when a script runs on a timer in a directory somebody else is also using. Then you ship a debug constant, a commented-out guard, a half-finished migration — something nobody wrote down, nobody reviewed, and which does not correspond to any commit you can go back and read.
Our first instinct was to point a five-minute watcher at the main checkout. That would have been a machine that publishes whatever is lying around, every five minutes, forever.
What we did instead
Connecting the repository is the correct fix, and we could not do it — the GitHub app was not authorised on a private repo and that needs a browser and a human. So the stand-in deploys from its own clone, which exists for nothing else:
cd ~/deploy-mirror
git fetch -q origin master
git reset -q --hard origin/master # safe here: this tree holds no work
railway up --service web --detach
The reset --hard is only acceptable because that directory can never contain anything a person cares about. Point the same three lines at a working checkout and you have written a tool that destroys uncommitted work on a timer.
Then wait for the thing you actually want
The last piece matters more than the deploy. railway up --detach returning zero means the upload was accepted. It does not mean the new code is serving.
So we put a build identifier in the app, returned from its health endpoint, and the deploy script does not consider itself finished until that string changes:
for i in $(seq 1 36); do
sleep 10
got=$(curl -s "$HEALTH?cb=$RANDOM" | grep -oE '"build":"[^"]+"' | cut -d'"' -f4)
[ "$got" = "$want" ] && { echo "LIVE: $want"; exit 0; }
done
echo "WARNING: uploaded but '$want' is not serving after 360s"
exit 1
First real run: detected the push, uploaded, and confirmed the new build id ninety seconds later. The value is not the ninety seconds. It is that a deploy which uploads and does not take now says so, out loud, instead of exiting zero.
The general shape
The failure here was not Railway's. It was that three separate things all report success without the outcome having happened: a push with nothing listening, an upload that is not a rollout, and a timer that runs on schedule while achieving nothing.
If you automate a deploy, make the last step check the running system for evidence the new code is there. Anything short of that is a machine that tells you it worked.
We build and run small products, and write up what breaks. Rebel Studios.
