Our deploy said Done, and the old code kept running
For weeks a deploy on one of our own projects warned, printed Done and restarted nothing, so the old version kept serving. What we found and fixed.
Somebody tells you a fix went live on Tuesday. On Friday you find out it was never there. On one of our own projects, every deploy for weeks ended with a warning about a missing privilege and then the word Done. The files were copied and the build was fresh. The application was never restarted, so the machine carried on serving the previous version, and the deploy log said it had succeeded.
That is our mistake, and the danger is the shape of it. A deploy that fails is an annoyance. A deploy that reports success without doing the thing it exists to do quietly separates what you believe is running from what is running. For a customer, the risk is being told a fix is live when it is not and relying on it.
What was actually going on
The deploy script runs as a dedicated user without special privileges, so that the code on the machine stays owned by that user rather than by the administrator. That was the right choice and we would make it again.
The last step restarts the application, and restarting needs a privilege that user does not have. The step tried, could not, warned and carried on. Another project on the same machine never hit this, because it has no dedicated user and runs the script as administrator, where the restart simply works. So the fault belonged to the project that had done the more careful thing.
A warning that does not change the outcome of the run is a lie told by the tooling. If a step is needed for the deploy to mean anything, it must fail the run. If it is optional, it should not be in the deploy.
The same session found the sibling problem. Images were being served with no caching instructions. The setting for it existed, was correct and was committed to the repository, in an example configuration file. The deploy copies the website files and never touches the web server’s settings, so a block that had been written and reviewed had never been live anywhere.
It also found two traps. The landing-page copy step deletes anything in the destination that is not in the source, so once a second application was installed inside that folder, the next landing-page deploy would have removed it. And the website’s production settings file was caught by a broad ignore rule, so it was never committed, and the server build fell back to development defaults: an address that points at each visitor’s own machine and an empty sign-in client id. No test catches that, because every test runs on a machine where that address happens to work.
What we changed
We gave that one user one permission: to restart one named service, using its full path and exact arguments, and nothing else. We checked the file with the validator before installing it, because the folder is shared by every project on the machine and one mistake would break the privilege system for all of them. We then proved the edges: the same user still cannot stop the service and still cannot get a shell.
We verified the effect, not the absence of the warning. The warning disappearing shows the grant was read. What shows the deploy works is that the service’s start time now matches the restart line in the deploy log.
The privilege file lives only on the machine, in no backup and no repository, so a rebuild would have silently recreated the original fault months later. It is now committed to the repository, with a step in the first-deploy instructions and a row in the runbook. The settings file was committed, since it holds no secrets, and the defaults now follow the build type, so a production build defaults to the production address.
What it did not fix
The same sweep found two leftovers from an earlier rewrite: a scheduled job still running every night against code that had been deleted, and a runbook that told a responder that data at rest is encrypted in a way it no longer is. Both were reported rather than touched, because they were not that day’s work. The caching setting and the deleting landing-page copy were found and described here; this write-up does not record either as fixed.
The pattern, for anyone who is told “it’s deployed”
Ask how they know the new version is running. A message saying Done proves nothing; a start time that matches the change does.
And after any fix applied directly on a server, ask what would happen if that machine were rebuilt from scratch tomorrow. Anything that would not come back has to go into the repository while you still know why it exists.
Where this ends up
That rebuild test is also what makes an engagement endable: no undocumented deployment steps, no credentials that live only in somebody’s head, and environment definitions that rebuild without the people who wrote them in the room.