An app that worked for us and failed for every new user
Days before store submission, a fresh API key got a 404 on a model the provider still listed. Availability is per account, and every earlier test used our own key.
Days before submitting a children’s learning app to a store, we tested it with a freshly created provider key instead of our own, and every generation failed with a 404.
The app generates quiz questions from a written lesson, and at that point there was no hand-authoring path left. A parent installing the app and creating their own provider key, as the setup flow tells them to, could generate nothing and so could publish nothing. The product was unusable for every new user and worked perfectly for us. That is how a release reported as tested ships broken: every test ever run had been an assertion about one account’s history. This was found before the submission, so no parent met it.
What was actually going on
The model identifier was pinned as a constant. The constant was correct. The model still existed and was still returned by the provider’s own list-models endpoint. It returned 404 for the new key, with a message saying it was no longer available to new users. Ours was an older key with the older entitlement.
Availability is a per-account fact, and the model list does not know whose account is asking. A list-models response tells you what the vendor has. It does not reliably tell you what your credential may call, whether the model accepts the request you are about to send, or whether it does what your feature needs. The same project hit all three failures on three different providers within a fortnight.
- Entitlement: the 404 above. Listed, real, and not callable by a new key.
- Modality: one provider’s integration discovered its model by taking the first entry the list returned. That entry had become a speech synthesis model, which returns 400 to every text call. The provider had been silently dead for an unknown length of time, and the failure looked like a generic bad request.
- Behaviour: on another provider, discovery could land on a reasoning model that spends its whole output budget on internal reasoning and returns nothing usable. Nothing errors. You get a successful response with no answer, the most expensive way to fail.
Two of the five configured providers turned out to be dead at the same time, one on the speech model and one on billing. A failover chain hides exactly this until someone goes and looks.
What we changed
Auto-discovery was replaced, because it hands the choice of what your product does to a list whose order and contents are controlled by someone else and change without notice. Selecting the first entry is a lottery run once per deployment.
The replacement is small. The model is pinned, with an explicit ordered fallback list, most specific first, ending in a rolling alias the provider maintains. The fallback advances on a 404 and on nothing else, since a rate limit is not “this model does not exist” and must not quietly change which model the product uses. Once a model has answered, the answer is cached. Five regression tests cover it, two of which were confirmed to fail with the fallback removed.
Adding a fifth provider, on a different API family, showed the other half: even with the right model chosen, the request shape is not portable. That family differed in five ways from what the existing code sent, and each difference is a 400 rather than a graceful degradation. The output token limit governs the model’s internal reasoning and its visible reply together, so a limit sized for the answer gives an empty answer once the model decides to think. The obvious control for thinking effort is unsupported on the chosen model and errors every call instead of being ignored.
What it did not fix
Nothing in the system ever stated which model each provider was resolving to. The speech-model failure went undetected because that answer was derivable at runtime and never derived, and a fact that is technically available and never surfaced is, in operation, a fact nobody has. It is still a gap to close: record where a person will read it which model each provider resolves to now.
An unsupported parameter that returns an error is better than one that is silently dropped, and both exist. The only way to know which you have is to send it once, deliberately, and read the response.
What to ask your own team or supplier
- When was the last time the integration was tested with a credential created today, on a new account, the way your users create theirs?
- Is the model pinned, with an explicit fallback list, or chosen by discovery?
- What exactly causes the fallback to move on? Is it one status code or any failure?
- Which model is each provider resolving to right now, and where can someone read that?
- If a provider stops working, who finds out, and how soon?
Where this ends up
Pinning the model, writing the fallback list down and recording what each provider actually resolves to is the unglamorous half of putting a model in front of a user. Which parts of that are specialist work and which apply to every engineer is set out under AI-first delivery.