Down for eight hours, and every light was green
The app was unreachable from 2 a.m. to 10:25. The cloud console said the server was running and healthy. The bill was unpaid and the notice went to one inbox.
Your field team starts its day and the app will not open. Not slow, not an error message: nothing answers. It stays that way through the morning. Somebody logs into the cloud console to see what has broken, and the console says the server is running, both health checks are passing, and the firewall is open. Everything is green. Nothing works.
That was 13 August. The outage ran from 01:54 to 10:25, about eight and a half hours, across every port on the production server. For Aries Agro, whose sales staff start shifts on their phones from eight in the morning, the last two and a half hours of it were felt by everybody.
What was actually going on
The usual suspects were ruled out first. The domain still pointed at the right address. The application had not been changed. The firewall rules were as they had been for months. The console genuinely reported the machine as running with its status checks passing.
The one signal that told the truth was the management agent, the small service on each server that reports back to the console. Its last successful check-in dated the cut to the minute, 01:54 on production and 01:53 on the test server, and the third server, on a different account arrangement, was untouched. Two servers on the same account cut off in the same minute is not a software fault. It was a billing suspension: an unpaid invoice, and the provider had withdrawn the network while leaving the machine “running”.
When the bill cleared, everything came back on its own. No restart, no redeploy. The portal answered, the API answered, the database was reachable on both servers, and the field team never knew what it had been.
The unpaid-bill notice had gone to exactly one place: the root email address of the account. Nobody had ever set the alternate contacts for billing, operations or security that the provider offers precisely so a notice like that reaches more than one person. There were no alarms and no notification topics on the account at all. The two cost alerts that did exist each went to a single person.
We found that out a month later, while checking how much the new AI assistant was costing, which is the kind of moment a gap like this tends to surface: not during the outage, but the next time somebody looks at the account for an unrelated reason.
What we changed
The three budget alerts and the daily cost-anomaly alert now reach three people, the client’s IT contact and two of ours, and the old recipients were kept. When the client’s contact changed, the new address was added before the old one was removed, so no alert was ever left with zero subscribers, even briefly.
The monthly budget was raised from 170 to 210 dollars, because August’s real spend was 197.78 and a budget that fires on a normal month is a budget everyone learns to ignore. A separate budget of 25 dollars a month was created for the AI service alone, with alerts at half and at full, so a runaway there surfaces on its own rather than hiding inside the total. For the record, the AI spend was 1.86 dollars for the month to date, 1.6 per cent of the bill, almost all of it from the two build days.
Every change was read back from the provider’s API after it was made, not assumed from the console.
What it did not fix
The alternate contacts are still not set. The provider’s API requires a phone number for each one, and we did not have the numbers to give it, so the single most important fix is waiting on a form. Until it is done, the next unpaid-bill notice goes to the same one inbox.
After the recovery, the management agent on both servers had not reconnected and needed a restart by hand. That agent is the fallback way into a server when the normal route is down, so for a while after the outage the fallback was down too. And the cloud credentials checked into two of the project’s configuration files turned out to be dead, rotated in an earlier clean-up, which cost time before the working ones were found.
The mechanism
A billing suspension withdraws the network and leaves the instance reporting healthy, so every dashboard built on the instance’s own view says green. The only thing that noticed was a service that has to reach out to somewhere else and stopped being able to. Outage detection that lives on the thing being detected is not detection.
Where this ends up
An account with one contact and no alarms is a normal state for a system that was built and then left running. Teams that build and run production systems of their own, which is what custom software development means here, treat the alerting as part of the delivery rather than something the client will get round to.