We retired a server and the bill did not move
A June clean-up was meant to cut the cloud bill. It stayed at about $198 a month: the saving was $4, not $52, and server backups had quietly stopped 27 days earlier.
In June you retired a server. It was a clean-up, and the point of it was that the cloud bill would fall. The July bill arrives and it is the same as before. So is August. You ask, reasonably, whether the clean-up was done at all.
It was. The retired server had cost 51.85 dollars a month and that line is gone. But in the same period a third server appeared for the website at 33.30, a third static address came with it, and 10.73 a month of data transfer that the old server had bundled into its price became metered: 198 gigabytes in August. The net saving was about 4 dollars a month, not 52. The bill had been flat at about 198 dollars since March, and the client, Aries Agro, deserved to hear that plainly.
We were only looking because we were costing an AI feature and needed to know what the account actually spent. The audit was read-only, six months of the provider’s cost reports set against the live inventory, and cost was taken from usage rather than estimated.
What was actually going on
Two things were leaking, and one thing had broken.
The test server ran 744 hours a month, every hour of every day, at 3.4 per cent processor use, for 43.54 dollars. It is used in office hours on weekdays.
The production server was being snapshotted every twelve hours with thirty copies kept, fifteen days of whole-server images, while the database was separately dumped every day to storage. The estate held 45 snapshots, 2,700 gigabytes provisioned, and the snapshot line had grown from 249 to 363 to 426 gigabyte-months over three months, 25.11 dollars in August.
The broken thing was more serious than either. The policy that takes those production snapshots was in an error state and had taken nothing since 12 August. That date is the billing suspension that took the app down for eight hours; the policy had failed during it and the service does not retry its way out of an error state. Its permissions read correctly. It simply sat there, 27 days with no server-level backup of production, while we paid for thirty stale ones.
Before touching anything we confirmed the daily database dump was healthy: that morning’s had landed at 02:00, 99.3 megabytes, thirty files under a thirty-day lifecycle. So the data was safe. It was the machine image that was not.
What we changed
A manual recovery point of the production volume was taken first. The script that deletes old snapshots refuses to run unless a completed snapshot newer than its cut-off exists, re-derives its target list from the live API rather than an earlier listing, and excludes anything that backs a bootable image. Then the backup policy was re-enabled, and retention retuned from every twelve hours keep thirty to daily keep seven, matching the website’s policy, which turned out to be healthy already.
The test server now starts at 09:00 and stops at 21:00, Monday to Friday, on a schedule whose role is trusted by the scheduler alone and scoped to that single instance. The Friday stop carries it through the weekend. The whole chain was verified with a throwaway schedule firing at an already-running instance rather than waiting until that night to find out.
Thirty stale policy snapshots and four old manual ones went: 45 snapshots and 2,700 gigabytes became 11 and 710. The owner then decided to deregister the three remaining bootable images and their snapshots too, after we confirmed nothing referenced them, leaving 8 snapshots and 340 gigabytes.
The test server’s new window is about 260 hours a month, roughly 28 dollars saved, and the shorter snapshot chains take a further 13 to 17 once they age in. With the images gone the saving lands nearer 45 to 50 dollars a month, taking the bill from about 198 to about 155. The earlier figures we had quoted were before tax and these are inclusive, which is the 1.18 difference.
What it did not fix
Stopping the test server takes its database, and anything polling it from the accounting side, down outside office hours. Storage and the static address are still billed while it is stopped, so the saving is on compute only. The test server is still a previous-generation instance the provider flagged in June, and no reserved pricing has ever been bought against about 107 dollars a month of steady compute; both remain on the list.
And the trade the owner accepted deserves stating: with the images gone, rebuilding a lost server means restoring a volume from a snapshot rather than launching an image. That is slower, and it was a decision, not a side effect.
The mechanism
Cost that moves between line items looks like no change. Snapshots that grow on a schedule and a test server that never sleeps are where a small estate leaks. And a backup policy that fails into an error state and stays there is indistinguishable, from the bill, from one that is working.
Where this ends up
Retiring a server is a small act of modernisation, and the lesson from this one is that a migration is not finished when the old machine is off. It is finished when the bill shows it and the backups are proven to be running.