Let's talk
ai

Asked for fifteen questions, given twelve, told nothing

A children's learning app returned twelve questions when a parent asked for fifteen and said nothing. Fail, truncate quietly, or return twelve and say so.

· · updated

A child's hands holding a tablet on a sofa showing the Sazinga Engage modules list.

A parent using a children’s learning app asked for fifteen quiz questions and got twelve, with no sign that anything was missing. When a system returns less than was asked for and reports success, every status report built on it is wrong by the same margin, and nobody can tell.

The app had three different answers to “how many questions can I ask for”. One screen offered a dropdown of four, six or eight. Another sent any number it liked to a server that silently clamped it to twelve. And none of them told the user what had happened.

What was actually going on

The clamp is the worst of the three. A parent who asks for fifteen and receives twelve, without a message, concludes one of two wrong things: that the model could only think of twelve, or that they mistyped. Neither is true. A limit exists in a layer they cannot see, and the system chose not to mention it.

When a producer you do not control cannot fully satisfy a request, there are three options, and the choice is a design decision.

  • Fail the whole request. Clean and honest, and it throws away everything that succeeded. For a generation that took twenty seconds across several calls and cost real tokens, discarding eleven good questions because the twelfth call was rate-limited is an expensive kind of purity.
  • Return less and say nothing. Cheapest to build. It turns an explainable shortfall into an unexplained one.
  • Return what you have and state the shortfall. More work, and the only option that leaves the user able to act.

There was a parsing case that mattered more than tidiness. A multiple-choice question is a list of options and a numeric index saying which is right. If the model returns a blank option, the tempting fix is to drop it. That shifts every later option down by one and leaves the index pointing at a different answer. The question then marks the wrong option correct and looks entirely normal. A child answers correctly and is told she is wrong, or the reverse, and the system pays or withholds points on that basis.

The write side had a separate problem. Adding a batch of accepted questions was a loop over a single-item endpoint, because no batch endpoint existed. Fifty questions meant fifty requests and fifty commits. A failure at question thirty left thirty saved and twenty not, and the screen had no way to say so.

What we changed

The third option is what the app does now. A large request is split into several provider calls of a size the providers reliably handle, and results accumulate. If a call mid-batch fails, the response carries the questions that were produced, the honest count and a note on what happened. Asking for fifteen and getting twelve now produces twelve questions and a sentence saying so. The parser that reads the model’s output used to discard the whole batch if one question was malformed. It now skips the unusable ones and keeps the rest.

A question with a blank option is rejected outright, not repaired. Repair only what is local, and reject anything that would need a reference elsewhere to be adjusted as well.

The write side became one endpoint taking the whole batch, validating every item before adding any, and committing once. The test for it was run against a deliberately per-item implementation first to confirm it fails there, since an atomicity test that passes against a non-atomic implementation is worthless.

Two smaller changes. A failed turn in the chat screen, where a parent talks to the model, now stays in the conversation marked as failed, still holding its text, with retry and edit, where before a timeout after twenty seconds also cost the paragraph the parent had typed. And while a generation runs the screen shows an indeterminate indicator, not a progress bar, because a made-up fraction reads as nearly done at the moment it has stalled.

What it did not fix

The asymmetry between reads and writes is a rule and not a cure. Partial generation is acceptable because the work is reproducible: ask again and you get more. Partial writes are not, because the user cannot tell which half landed and re-running may duplicate. The policy fits this app. It does not remove the chance that a model returns fewer items than requested, only what the user is told when it does.

What to ask your own team or supplier

  • When a system returns less than was requested, where is that stated to the user and in the response?
  • Does the API have a name for a partial result, or is it dressed up as success or failure?
  • For every operation that can half-complete, can it safely be repeated? If so, keep what succeeded; if not, all or nothing.
  • When generated data is cleaned up, can any step move items in a list and invalidate an index stored elsewhere?
  • Does an error path ever throw away something the user typed?

Where this ends up

Deciding what a model is allowed to return, and what the screen says when it returns less than that, is guardrail work between a model and a person. Where it sits in an engagement, and where model assistance does not help at all, is under AI-first delivery.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch