Let's talk
engineering

The app said no connection. Our own patience ran out

Users were told to check their wifi. The network was fine. Our own timeouts were shorter than the work, and a generated answer was paid for and thrown away.

·

A kitchen table by a rainy window with cups of tea, coloured pencils and blank drawing paper.

Someone opens your app, presses the button that asks the AI to draft something, waits, and is told there is no connection. They know there is. They are on the same wifi as everything else, and the key they are using for the service is a good one.

That is what happened in a children’s learning app, where a parent uses AI to draft a lesson. The message was wrong, and so was the blame. The connection was fine. The app had simply given up before the answer arrived, and by then the answer had already been generated and paid for, and was thrown away.

A second fault of the same kind sat beside it. The first person to sign in after the server restarted got a failure, and the sign-in screen again told them to check their wifi. Nobody needs a support call to find out that a message like that is a lie.

What was actually going on

Both were timeouts that we had set ourselves.

The app cut off every request after 20 seconds. Behind the button, the server tries the AI providers in turn, and it pays for each one that fails before the one that answers. Measured on the real keys, that was 0.7 seconds on the first, 5.6 on the second and 15.2 on the third. That adds up to 21.6 seconds against a budget of 20.0. The third provider answered, the app had already walked away, and the result was discarded.

The sign-in fault was the same shape. The server checks a sign-in against Google’s published signing keys, and fetches them once after each start. On the first fetch that took 11.7 seconds, and the server’s own limit was 10. So it failed, and the first person through after every restart saw an error.

What we changed

The four routes that call an AI provider now allow 120 seconds. We made the timeout a kind of network error inside the app, so nothing that already handled network errors had to change, and we stopped the app claiming the network was down when the delay was ours.

For sign-in, the limit went to 20 seconds with one retry, and if the keys are already held from earlier, a slightly old copy is used instead of failing everyone during a hiccup. The sign-in screen now gives three different messages, so a server fault no longer says “check your wifi”.

What it did not fix

A 120-second limit means a person may wait up to two minutes before hearing that something is wrong. We did not shorten the chain of providers, and the log does not record any change to how long the slowest path takes. We only stopped giving up on it.

Nor did we remove the cost of trying several providers. The server still pays for each one that fails before the one that works. The fix made sure the answer, once paid for, arrives.

The pattern, for anyone whose app calls slow services

Any limit you set on the waiting side must be longer than the slowest normal path behind it, including every fallback. If the server may try three services in turn, the screen cannot give up after the time of one.

Then look at the wording of your error messages. “No connection” is a claim. If your own system is the one that stopped waiting, say that. A wrong message sends a person to fix their wifi, and sends you nowhere.

You can check this in your own business without any code. Ask how long the slowest normal request takes, and compare it with how long the screen waits. If the second number is not clearly larger than the first, some of the failures you hear about are your own.

Where this ends up

This is the kind of fault that turns up when an app calls AI services and has to live with their timing, which is the work described on our AI and data assistants page.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch