Full marks on the safety quiz, and nobody read the lesson
Full marks on a safety quiz proves nothing about whether the lesson was read: the careful-sounding answer is obvious. What the app could and could not record.
You issued the safety lesson, the quiz came back with full marks, and you are fairly sure the lesson itself was never opened. The score is the only evidence you have, and it is telling you the opposite of what you suspect.
This one was a parent, not an employer, and the lesson was for a ten-year-old on a learning app we build and run. But the question is the one every business asks of its training the week before an audit: did they read it, or did they go straight to the quiz? The parent asked us, and the honest answer was that nobody knew and the app had never been built to know.
What that costs depends on the subject. On arithmetic, a skipped lesson costs a little learning. On safety — the one subject the lesson exists for — a full score with the reading skipped is worse than no score, because it puts a tick against a name and the tick means nothing. Anyone who has signed off a safety refresher on the strength of a quiz result is holding the same tick.
What was actually going on
The app has a reading screen with the written lesson on it, and a quiz attached that pays points. Every record the app keeps about a child’s progress describes the quiz: the score, the attempts, when it was last taken. Nothing anywhere records that a lesson was opened. The step that starts a quiz checks that the lesson was assigned to this child and that it is published, and does not check — because it cannot — whether the lesson was ever displayed.
So the only available evidence was the quiz score, and the obvious argument is: if she scored well, she must have read it. That argument is wrong in a specific way that only shows when you look at it subject by subject rather than on average.
On a maths lesson it holds. You cannot guess seven hundred and twenty divided by eight from a list of four numbers; the score is evidence of something. On a safety lesson it collapses. Every question of the form “what should you do if…” has one option that sounds like the careful, sensible answer, and a ten-year-old picks it out of the list without having read a word. She could score full marks on a subject she never opened.
So the score is strongest exactly where it matters least, and useless where the whole purpose of the lesson is that something was absorbed before it is needed. A stand-in measure’s accuracy is not one number. It is different on every part of what it measures, and the useful question is where it is worst, not what it is on average.
What we found one layer down
The number we had spent an hour concluding did not exist was being calculated the whole time. The reading screen draws a progress bar as you scroll, and puts you back where you were when you return to a half-read lesson. To do that it works out, continuously, how far down the lesson you have got — and keeps that figure on the tablet, and never sends it anywhere. It had been built to draw a bar, not to record a fact about the child, so it was written to the device and thrown away.
That is a pattern we now look for deliberately. Before building a way to capture something, search for the places the system already works it out for its own small purpose: scroll positions, “last seen” markers, retry counts. They are usually accurate, computed close to the event, and discarded. Turning one into a recorded fact is far cheaper than measuring from scratch, and more reliable, because something the user can see already depends on it.
What was rejected, and why
The cheapest option is a minimum time on the page: ninety seconds on the reading screen counts as read. Rejected on two grounds. A child works out within a day that leaving the tablet open and walking away passes it, and that cannot be told apart from reading. And it punishes a fast reader who genuinely finished and is told she has not. A measure that a cheat and an honest fast reader both fail is not measuring reading.
The scroll position is better and still weak. It shows the words passed under the eyes. It shows nothing about attention.
The most useful change was not a measurement at all. It was writing questions that cannot be answered without having done the thing. A later set of lessons built around an exercise asks what ended up in a particular cup after a swap, how many times a repeated pattern occurred in a physical layout, where an item lands under a rule applied top to bottom. Those answers exist nowhere except in the activity. They are a weak but real check, and they cost nothing to build, because they are just how the questions are worded.
What it does not fix
None of this verifies understanding. At best it verifies exposure. For a safety subject that distinction is the whole thing, and no amount of engineering closes it. The real check is a person having a conversation — a parent at the kitchen table, a supervisor on the floor — and the most the software could do was write the lessons so they give that person an opening for one.
We record that as a limit rather than a plan, because it is the correct conclusion and not a gap. The instinct when a product cannot measure something is to build a measurement. Sometimes the right answer is that the measurement is not available to software, and the product’s job is to make the human check easier rather than to stand in for it.
The mechanism, briefly
The progress record is keyed on the child and the quiz, and every column on it describes quiz performance. The child’s device has four routes it can call, and none of them is a mark-as-read. The reading screen computes a live scroll fraction for its own progress bar and persists it in local storage on the device only. Two things carry forward from this. Search for the signal you already compute and discard before you build a new one. And when you accept a proxy, name the part of its domain where it fails before you ship it — because if you do not, you will discover it in the one subject where being wrong actually costs something.
Where this ends up
The same question, asked of a workforce instead of a child, is what Sazinga Engage is built around: comprehension checks generated from the document you actually issued, so that answering them correctly demonstrates having read that document rather than knowing the sensible-sounding answer.