Let's talk
engineering

The portal went down for a few seconds, many times a day

Staff said the server was down. It was healthy and answering in 40 milliseconds, yet the office saw 219 error pages in a day. The cause was a customer's browser.

· · updated

A laptop at a hire-office desk showing the availability calendar in Sazinga Rentals.

Your office staff tell you the booking portal is down. You try it and it works. Ten minutes later somebody else gets an error page, and then it works again. Nobody can say when it will happen next, only that it happens in bursts, several times a morning, and that whatever they were typing when it happened is lost.

For a self-drive car rental operator in Goa (the system behind GOI Car & Taxi) this was the state of the back-office portal for at least four days in August. Counted from the logs: 219 error pages on one day, 228 the next, then 86, then 20. The machine was not busy. It had plenty of memory and disk, the database was idle, and a request sent straight to it came back in 40 milliseconds. The production system had restarted itself 340 times; the identical test system had restarted 80 times. Something was killing the live one and not the copy.

What was actually going on

The portal has two kinds of visitor. Staff sign in with a password. Customers open their own booking page with a booking code, and the portal gives them a short-lived pass so they do not have to type the code again. Both passes were kept in the browser under the same name.

So a member of staff who had, at any point, looked at a customer’s booking page in the same browser was carrying the customer’s pass. The next time that browser asked the staff side for anything, it handed over the wrong pass. That should produce a polite “not authorised” and a redirect to the sign-in screen. Instead, because of where in the code the refusal was raised, the whole server process fell over. A supervisor restarted it within two to five seconds, and every request from anyone in that gap got an error page. The log held 315 such crashes.

The test system never showed it because no customer had ever used its booking page, so no browser there had ever carried the wrong pass. The fault was invisible in exactly the place it was supposed to be found.

Two other things turned up in the same look. A handful of brief edge errors on both hostnames were between the content network and the server, not ours, and cleared on retry. And a scanner was hammering the older public booking site looking for a common web-shell, 584 of that day’s 604 error lines, without exhausting anything.

What we changed

Both places that refuse a bad pass now hand the refusal back to the framework instead of throwing it from inside a callback the framework cannot see, so a bad pass gets a 401 and the process stays up. Narrowing the trigger while testing mattered: an expired or malformed staff pass had never crashed anything, because it was rejected earlier on a path the framework does catch. Only a customer pass reaching a staff endpoint took the process down.

The second finding was that the portal’s error handler, the part that turns a fault into a proper response and a log line, had been written and imported but never switched on. That was why every refusal had come back as a generic error page rather than a sign-in redirect. Switching it on nearly caused a worse problem, caught in a pre-deploy recheck: as written, it would have logged the full request body on a failed sign-in, which is the email and the typed password, for every wrong password attempt. There were 145 of those in the current log. The handler now redacts passwords, one-time codes and tokens from both headers and bodies before anything reaches disk.

Verified by mounting the real code on a live server and firing six kinds of pass at it: no pass, garbage, expired, customer, wrong secret, all returned 401 with the process alive; a valid pass returned 200. The same six against a verbatim copy of the old code killed it on the customer case. Deployed. The exact request that had been taking production down now returns 401 in 16 milliseconds with the process untouched.

What it did not fix

The same defect class, a fault raised from inside code the framework cannot catch, exists in four more places in an older set of order routes. Two weeks of access logs showed zero requests to them, so they were left out of this deploy to keep it small and flagged for follow-up. They are still there.

The browser still stores the customer’s pass and the staff pass under the same name. The server no longer dies when it receives the wrong one, but the confusion that sends it has not been removed.

The mechanism, briefly

The auth middleware threw from inside passport’s custom callback, which the strategy invokes asynchronously via strategy.fail(), outside Express’s try/catch; Node treats that as an uncaught exception and exits. Replaced with next(err), and the error handler that app.ts imported but never registered was registered last, with header and body redaction and a one-line log for sub-500 responses so routine 401s stop burying real faults.

Where this ends up

An office that sees error pages in bursts stops trusting the screen, which is why Sazinga Rentals treats a refused request as a normal answer the server gives, not an event it survives.

This came out of building Sazinga Rentals

Bookings, availability and the fleet standing behind them. The problem above is one we met while building it, and what we did about it is in the product.

If you run something like this, there is one thing you can do without a call: send one week's booking sheet.