Which pass of the press is losing you product?
Yield was recorded per batch, so a bad pass through the press was hidden. Recording output per run fixed that, but only from a date forward.
You can see how much oil a whole batch of seed produced. You cannot see which pass of the press is losing you product. That is the number that would tell you whether the problem is a machine setting, a temperature or an operator, and a batch total cannot give it, however carefully it is added up.
That is the question the person who asked for this change was actually asking. They were not asking for a tidier database. They were asking to see which pass was losing them product.
What was actually going on
Production works at two levels. A batch is a planned quantity of one raw material to be processed. A run is one pass through the press. A batch contains several runs, because the machine takes a fixed charge and a batch is bigger than one.
Each run recorded how much raw material it used. The output, the amount of product that came out, was recorded against the batch. Yield is output divided by input, so with input on the run and output on the batch you can work out the yield of a batch and never the yield of a run. Batch yield tells you what a lot of material produced, which is useful for costing and useless for operations. Run yield tells you that the third pass of the day is worse than the first two.
I found the history by reading the database itself. The table for output still carried its old names, because renaming a table leaves its internal labels behind, and nobody tidies them. An insert failed and the error message named a table that does not appear anywhere in the current system. Next to it were two columns pointing up the hierarchy: the original batch column, still required, and a newer run column, optional.
Read together, they tell the story. Output used to be recorded against the batch. Someone needed it against the run. The new column could not be made mandatory, because existing rows had no run to point at, so it went in as optional, the old one stayed, and the code started writing both. That is the correct way to do it. It leaves a permanent ambiguity: for rows older than the change, the run-level question has no answer.
What we changed
Output is now recorded against the run as well as the batch, so yield can be worked out for each pass, which is the level at which someone decides what to adjust. When a ratio matters, both sides must be recorded at the level where the decision is taken. Both designs store true numbers, so nothing in the database looks wrong. One of them simply cannot answer the question anybody wanted to ask.
The work also turned up a small duplicate: two identical safeguards on the same link, added under two different naming styles by two people, or one person twice. The database accepted both. It costs a little and does no harm, but it is the cheapest sign that two people worked on the same thing without seeing each other’s work.
What it did not fix
The old rows were never filled in. Nothing anywhere records which run produced which output, so the data to rebuild it does not exist. The run-level yield report therefore starts on a date, and says so on its face, which is the only honest way to present a number that cannot be worked out for the earlier period.
The pattern, for anyone who wants yield or cost per step
Check the level at which each side of your key ratio is recorded. If output is logged per batch and input per pass, you will only ever see the average. When a record moves from one parent to another, write down the date the new way began, and put it where the next person will see it. The empty fields will outlast everyone who remembers what the emptiness means, and a report that treats empty as zero is wrong for every row before that date.
Where this ends up
That schema is the one behind Sazinga Factory, where output is recorded against each run as well as each batch, and a figure the older rows cannot support is shown as starting on a date rather than quietly counted as zero.