Our backups stopped for four days and nobody was told
A nightly database backup on our own server wrote nothing for four days, from four separate causes, with no alert. A dead certificate renewal hid the same way.
On one of our own servers, the nightly database backup produced nothing for four days. Nobody knew, because a job that fails quietly and a job that succeeds look identical from every direction except the one nobody checks.
That is our mistake, and it is the kind a supplier should own. Had that database been lost in those four days, there would have been no recent copy to restore from. It is the risk a customer carries whenever the people running their system cannot prove the last backup exists. Here it was our own system, and we had not proved it either.
What was actually going on
When we looked there were four independent causes stacked on one another, plus a fifth problem with the schedule.
The backup command was stopping to ask for a password, and running unattended it failed at once. That alone would have been the whole story. But the server also had no permission to write to the storage area, so even a finished backup could not have been saved, and that fault was hidden behind the first: you cannot discover an upload problem while nothing is being uploaded. The rule meant to expire old backups had never been applied, which would have shown up later as a storage bill. And a cleanup fault left roughly 150 MB of waste on the disk every day. On top of that the schedule was set in the wrong time zone. It was meant to run at two in the morning and ran at half past seven local time, in the middle of the working day.
Only one of these could be seen at a time. Fix the password and you find the permission. A job that has never worked end to end does not have a bug; it has a queue of bugs, and you only ever see the first.
We found the same shape elsewhere. Automatic renewal of security certificates had not run since early June. The scheduler entry that drives it had been switched off in a way that meant it could never start. Next to it sat an older scheduled entry that looked exactly like a safety net, but it switches itself off when the newer scheduler is present, which it was. So renewal was off while a plausible-looking entry suggested it was covered.
We found it ten days before four certificates would have expired, which is the point at which visitors to those sites would have begun to see security warnings. Switching the scheduler back on renewed all four in a single run.
What we changed
We fixed the four causes and the schedule, and we turned the renewal scheduler back on.
We also changed how we watch. A checker that runs on the server it is watching cannot report that the server is down. During a real outage the connection answered in under half a second and the name lookup worked perfectly; only a full request to the application showed the error. So checks now look for the right status and a known piece of text on the page, because a success code carrying an error page should count as down, and a response taking ten seconds against a normal one second counts as down too.
What it did not fix
The write-up named the top remaining gap in one sentence: the job failed silently for four days. The four faults were ordinary and quick to fix once noticed. The failure was that nothing said anything for four days. An alert for silence, meaning no success signal inside the expected window, is the piece this account does not claim is built.
The pattern, for anyone who relies on a backup
A scheduled job needs three things: it runs, it reports failure, and it reports silence. Most have only the first. A job that fails loudly is a solved problem. A job that never runs at all produces nothing, and nothing is exactly what a healthy quiet system also produces.
Ask this: if this job silently stopped existing, how long until somebody found out? If the answer is “when we need it”, you do not have a backup, you have an intention.
Then go and look at the most recent backup of a system you are responsible for. Check its time and its size. If you cannot do that in under a minute, that is the finding.
Where this ends up
Verifying a system by what it produces, rather than by what it is configured to do, is how we run an ODC engagement, where what you inherit is something that can be proved to work without the people who built it in the room.