Booking site down for hours with hardly any visitors
A booking site went down for four hours with about twenty requests a minute and an idle database. A bigger server would have made it worse.
Your booking website stops answering. Every page times out. You restart it and it works for a few minutes, then stops again. The obvious reading is that you have outgrown the server, and the obvious remedy is to pay for a bigger one.
In this case that would have made things worse. The site stayed down for four hours, and while it was down almost nobody was using it. Every hour it was dead was an hour in which a customer could not make a booking.
What was actually going on
The measurements were clear. Requests were arriving at roughly twenty a minute. The server’s load average was 0.15, which is close to idle. The database was doing nothing at all: every connection was asleep and no query was running. Nothing was working hard. Nothing was working.
What the application’s workers were doing was waiting, each one holding a connection to the site’s own public address.
The front of the application fetched its data from the back of the same application, but it did so by going out to the internet and back in through the public web address, instead of calling it directly. So every page view used two workers: one for the visitor’s request, and one for the request the site made to itself.
The site was set up with five workers. Five visitors at the same moment took all five, and each was waiting on a second request that needed a worker, and none was free. Nothing could finish, so nothing ever freed up. It could not recover on its own.
Two other faults turned this from a stall into a four-hour outage. The internal call had its timeout set to zero, which in that software means wait for ever. And the site already had a good fallback, to show a saved copy of the data if the call failed. That fallback never ran, because the call never failed. It just hung. Two workers had been stuck for fifteen to eighteen hours before anyone noticed, so it took only three more visitors to finish the job.
Any one of the three faults, fixed on its own, would have stopped the deadlock: the site calling itself, the unlimited wait, the small pool of workers.
What we changed
We fixed the faults in the internal call, so it no longer hangs and a stalled call can fail instead of waiting for ever, which lets the saved-copy fallback do its job. After the fix the internal call answered in 0.04 seconds, and ten simultaneous requests all completed in about half a second. Five had been enough to lock it.
What it did not fix
Restarting cleared the problem for a few minutes each time, and nothing in a restart addresses the cause. Adding workers would have done the same: it would have raised the number of simultaneous visitors needed to lock the site from five to ten, postponed the next outage and made it certain.
There was also a warning. The worker pool logged that it had reached its maximum about a minute before the first timeout. It is the kind of line that appears in the log during ordinary busy moments, which is why nobody acted on it.
Why a site ends up calling itself
This is not an unusual mistake. The application needs data from its own back end, and that back end has a public address already written into the configuration, because the browser needs it. Using the same address from the server side is one line, and it works in development and in testing. It works in production too, right up until the number of simultaneous visitors reaches the size of the worker pool. The correct forms are to call the internal function directly, or to use an address that never leaves the machine.
The pattern, for anyone whose site goes down
Before buying capacity, find out whether the system is busy or blocked. They look identical from the outside and need opposite remedies. Three questions separate them:
- Is the database busy? An idle database means the application is not asking it for anything.
- Is the server’s load high? A blocked server is quiet, because waiting uses no processing power.
- What are the workers waiting for? Here they were waiting on the machine’s own address, which was the whole diagnosis.
If a site is unresponsive and all three say idle, more capacity will not help, and you should ask whoever runs it to find what the workers are waiting on.
Where this ends up
Finding the cause of a fault in a running system, before changing or replacing anything, is the first step of enterprise software modernisation, and it is usually where the most useful evidence turns up.