Our AI chat showed a parent a line of garbage characters
A reply from an AI assistant was cut off mid-sentence and its hidden control text reached the screen, with a button that would have put it into a child's lesson.
A parent using the AI chat in a children’s learning app saw a line of at-signs and capital letters in the middle of the conversation. That alone looks broken. The worse part was the button underneath it, which offered to accept the reply as a lesson. One tap would have written those characters into the body of a lesson that a child then reads.
It is the kind of fault an AI feature can ship with unnoticed: it only happens on the longest, most valuable requests, which are the ones nobody tests with.
What was actually going on
The chat lets the AI model mix two things in one reply: ordinary conversation, and an offered draft of a lesson. To tell them apart, the model wraps the draft in a distinctive opening string and a distinctive closing string, and the app looks for text between the two.
The model has a limit on how much it may write in one reply. When a reply hit that limit in the middle of a draft, the opening string was there and the closing string never arrived. The app found no complete draft, decided the whole thing was ordinary chat, and displayed it exactly as received, control text included.
Truncation is not bad luck. With a length-limited generator it is a routine outcome, and it is most likely on exactly the requests where the content is longest.
The fallback made it worse. When no draft was found, the accept button used the whole reply as the lesson. It sounds like a sensible safety net: if we cannot find the draft, use everything. In practice it sent unchecked model output straight into something a child reads.
The cause of the truncation was more embarrassing. The chat’s length limit was smaller than the “long” lesson length the same product offers elsewhere, when a parent asks for a long lesson. The chat is the one place with no length control, so it is the place most likely to be asked for something long, and it had the smallest budget in the system. There were three limits in three places, set at three different times, and nobody had ever put them side by side.
What we changed
A reply with an opening marker and no closing one is now read for what it is. Everything after the opening is treated as the draft, and the result is flagged as cut short so the screen says so instead of pretending the lesson is complete.
The accept button no longer falls back to the raw reply. A reply that cannot be read as a draft is reported as unreadable, and the raw text is shown to the parent to look at. A fault now degrades the feature (no draft is offered) rather than the content (broken text published).
The chat’s length limit was raised well above the longest lesson preset, and the instruction given to the model now tells it that a budget exists, so it plans a shorter piece rather than being cut off.
One related decision is worth recording. The conversation history is capped, so something has to be dropped. Dropping the oldest message is wrong here, because the oldest message is the parent’s original instruction: what the lesson is about, who it is for, what tone. So the first message is always kept, recent messages are kept as far as they fit, and the middle goes. The model keeps the brief and the latest exchange, and loses the back-and-forth of “shorter please”.
The quiz generator avoids the whole class of fault. Where the model provider supports a structured answer format, the quiz path uses it, so the model cannot return something unreadable. There are no markers on that path, so there is nothing to be cut in half.
What it did not fix
A cut-off reply is still cut off. The parent sees the flag and a partial lesson, not a finished one. The length limits are now consistent, but they are still limits.
The record kept of each generation stores counts, the model used and the provider, and never the message text. The cost is that we cannot look at what a parent asked when they report a bad result. That is the right trade in a product where the messages are a parent writing about their own child, but it does make diagnosis slower.
The pattern, for anyone adding an AI feature to a product
Ask three things about every place the AI writes into your product. What does a customer see when its answer is cut in half? Is there any path on which unchecked AI text can end up in a field other people read, even as a fallback? And where are all the length limits written down, side by side?
If the third answer is “in three places, set by different people”, the smallest one is governing your product, in the place you least expect.
Where this ends up
Deciding what happens to an AI reply that arrives in pieces, and what a screen may show when it does, is guardrail work rather than prompt work. Where that sits in an engagement, and where AI assistance does not help at all, is set out under AI-first delivery.