The AI feature failed half the time and blamed the network
A question generator failed on about half of requests and blamed the network. The network was fine: the system had stopped after one of its five AI suppliers.
You have a feature in your app that takes a document — a safety brief, a procedure, a page from a handbook — and writes questions from it, so you can check that the people you issued it to have actually read it. About half the time it does not work. The screen says there is no connection. Your staff check the Wi-Fi, the Wi-Fi is fine, they try again, and it may or may not work the second time.
That was the state of a question generator we look after. The feature is the reason people open the app, and a feature that works one time in two is one that people stop opening. It was also costing money on every failure: the work of writing the questions had usually been done, and paid for, by the time the phone gave up and threw the answer away. And because the message blamed the network, everyone who tried to help looked at connectivity, which is precisely where the problem was not.
What was actually going on
The system does not rely on one supplier of AI answers. It has five, each with its own account, arranged in a queue: if the first is busy, out of credit or down, the request moves to the next. With five suppliers, an outright failure should have been rare.
The first guess was the accounts — a card had expired, a free allowance had run out. Reasonable, and wrong. The second guess was that the answers were coming back in a form the app could not read. Also reasonable, also wrong.
What settled it was three lines of the system’s own log, next to each other. The first supplier said “payment required”. The second said “success”. The app then reported “empty reply”. The second supplier had answered successfully, and the answer was blank. The queue treated a success as a success and stopped. Suppliers three, four and five — three working accounts with credit on them — were never asked.
Two more problems were hiding behind that one, and only appeared once the queue actually moved on.
The supplier whose account had lapsed was still first in the queue on every single request. There was a way to test an account, and the test correctly reported it as failing, but nothing wrote that result down anywhere, and the rule that switched off bad accounts covered “wrong password” and “not allowed” but not “unpaid”. So every request wasted two round trips on a dead account before anything useful happened.
And the phone was giving up too early. It allowed 20 seconds for any request, then reported the wait as a network failure. Measured on a real request: the first supplier was skipped, the second answered in 0.7 seconds with something unusable, the third took 5.6 seconds to say it was too busy, and the fourth answered properly at 15.2 seconds. Total, 21.6 seconds — against a phone that stopped listening at 20.0. The work was done. The answer existed. The user was told there was no connection.
What we changed
The queue now moves on when an answer is unusable, not only when a supplier reports a fault. A blank success is treated exactly like a busy signal: try the next one. For the question generator there is a stricter test still — an answer that is not blank but is not a set of questions counts as no answer.
An account’s health is now recorded. When it is tested, the result is kept: when, whether it passed, and what went wrong. An unpaid account is switched off with the reason stated, rather than sitting at the front of the queue.
And the phone’s patience now depends on what it is waiting for. The handful of screens that wait on an outside supplier get a long allowance; everything else keeps the short one. When the app does give up waiting, it says that it gave up waiting, which is a different message from “no connection” and sends whoever is helping in the right direction.
What it did not fix
None of this makes any single supplier more reliable, and a user on the question screen can now wait longer than before — the honest alternative to being told a lie at twenty seconds. The suppliers still fail; what changed is that a failure at one of them no longer ends the request.
The mechanism
For anyone who runs a system with a fallback chain, the shape of the fault is worth knowing, because it is common and it does not look like a bug. Nothing crashed. Nothing was mistyped.
The retry loop advanced on transport-level failures — a rate limit, a server error, a dropped connection — and treated anything else as a result. The check for an empty reply ran after the loop had returned. Each piece was correct on its own. Arranged that way, the check could never inform the retry, so a provider returning a successful nothing terminated a chain of four working alternatives. The fix was to raise the “unusable reply” error inside the loop, so it advances the rotation the way a rate limit does, with an extra predicate on the quiz route because prose where structured questions were required is the same failure as no reply.
The key store’s definition of a bad key covered 401 and 403 and not 402, and the test endpoint threw its result away. Persisting the outcome of each test, and treating a billing failure as its own class with its own stated reason, is what keeps a dead key out of the rotation.
The client applied one flat 20-second timeout and mapped an abort to a network error. It now has a per-route budget and a distinct timeout error type, so “we gave up waiting” is never reported as “we could not reach the server”.
All three share a structure: a component with a local definition of failure narrower than the whole system’s. The rotation’s did not include a valid-looking empty response; the key store’s did not include an account out of credit; the client’s did not account for its own retry chain running behind the request. Seams like these do not show up in unit tests, because unit tests exercise components. They show up in production as an intermittent failure rate nobody can reproduce — which is exactly how this one presented. The transferable rule is to decide what “success” means once, and make sure every layer that can retry is using that definition rather than its own.
Where this ends up
The generator this happened in is the same one Sazinga Engage uses to turn the safety brief or procedure you already have into comprehension checks. A check that arrives one time in two is not a check, which is why “did a usable question come back” is now the only definition of success the whole chain uses.