Let's talk
ai

The assistant invented four requests the child never made

A parent asked an assistant what their child had requested. It listed four things. The child had asked for one reward, four times. The prompt already forbade it.

·

A child's hands holding a tablet showing the Sazinga Engage dashboard, seen from a sofa at home.

You have an assistant that sits over your own records and answers questions in plain words. In this case the records belong to a learning app for children: a child completes lessons, earns points, proposes goals and asks for rewards, and a parent approves or declines. You ask it a plain question: what has the child asked for? It answers with a unicorn plush, a set of markers, a book and a movie night. Confident, tidy, four items.

The child had asked for one reward, four times.

A parent catches that, because a parent knows their child. Swap the household for a business and ask the same shape of question, what has this customer ordered, which suppliers are waiting on us, and you would not necessarily know. The answer would look exactly as tidy. That is the cost here: not four wrong items on one screen, but an answer that could not be told apart from a right one by anybody who did not already know the truth.

This was found on the day the assistant went into production, by the owner testing it. It was the last of five faults found that day, and the only one worth writing about, because the other four were plumbing and this one was the model doing exactly what a model does.

What was actually going on

The assistant keeps an audit trail of every turn: which lookup it ran, what came back, how long it took. That trail named the cause without any guesswork. The turn that invented the four items had run no lookup at all. It had answered in one second, from the count in its own previous sentence. An earlier answer had said, correctly, that there were four requests, because that answer had gone to the records. The follow-up, asked what they were, did not go back. It took “four” from its own prose and filled in four plausible things a child might want.

The reason is structural. When a conversation continues, the model is shown the words of the earlier turns and never the rows behind them. It saw a sentence saying four, and the sentence was the only thing it had. The written instructions to the assistant already said, in as many words, not to invent names. That instruction was in front of the model when it invented four. A prompt is a request, not a rule.

What we changed

The rule is now enforced by the system rather than asked of the model. If a turn ran no lookup and retrieved no reference passage, then nothing grounded the answer, and the answer is not sent. The model is asked once more, told explicitly that an earlier turn is not a source, and if it still will not go and look, the assistant refuses. The refusal is recorded as a refusal, so refusals can be counted rather than being disguised as answers.

Two smaller faults were visible in the same screenshot and were fixed in the same pass. A question about today had been answered from the child’s whole assigned list, because the lookup behind it carried no dates. And the assistant’s formatting was reaching the parent as literal asterisks, because the phone app shows text exactly as sent. Both were fixed on the server rather than in the app, which matters, because the app had already been submitted to the store and could not be changed that day.

What it did not fix

A grounded answer is not a verified one. The check confirms that the model looked at the records before answering; it does not confirm that what it said matches what it found. A model can read four rows and still describe them wrongly, and this rule would not catch that.

A refusal is also a non-answer. The parent who asked gets told the assistant would not answer, which is better than four invented items and worse than the truth.

And it was found by a person who knew the truth. There is no automatic check that an answer agrees with the rows, and the cases that matter most in a business, the ones nobody already knows the answer to, are exactly the ones such a check would have to cover. The same day’s work recorded two gaps in the app itself that are unrelated and still open: a fully funded goal never moves to “reached”, and a task can ask for photo evidence that the child’s app never collects.

The mechanism, in plain words

Each assistant turn either calls a tool, which runs a scoped query against the household’s rows, or retrieves a passage from the reference corpus, or does neither. The audit row for each turn records which. The grounding check reads that row: no tool call and no retrieval means the reply is rejected before it reaches the parent, the model gets one further attempt with an explicit instruction that conversation history is not a source, and a second ungrounded reply is recorded with outcome=refused. Conversation history replays prose; it never replays result sets, so a follow-up question is a new question and must run its own lookup.

Where this ends up

The same principle runs through our AI data assistant work for businesses: the model chooses what to look up, the database finds the rows, and an answer that did not come from the rows is not an answer. A children’s learning app is where we watched it fail in one second with four invented items, and the fix is the same whatever the records are.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch