Let's talk
operations

A build report said passing; the build had failed

A mobile release build went through a pager and the report carried the pager's exit code. The build had failed and the status report said it had succeeded.

· · updated

A hand pressing a blue stamp onto a blank sheet of paper on a wooden desk.

How do you know your supplier’s report that all tests pass is true? In this case, the honest answer is that I reported a failed build as passing, and it was my own report.

A mobile release build produces a great deal of output, so I piped it through a command that shows only the last part. That command succeeded, and I read it as the build succeeding and said so. In a shell pipeline the exit status is the status of the last command, and the pager exits zero whatever it is given. The build had failed, several hundred lines above what I was looking at.

I corrected it in the same session, which is the only reason it did not become a fact somebody else planned around. That is the cost to you: a wrong “done” turns into a decision made on a false fact.

What was actually going on

The thing measured and the thing reported were different, and nothing in between complained. The same class appeared from every direction.

One round reported all twenty-five test suites green. Re-running them produced twenty-two green and three failures. The report was not dishonest: it described a run that had happened. It was simply not a verification, because nobody had checked it against the tree as it stood.

Another round left two lint errors in the codebase and finished without producing a report at all. Absence of a report reads as absence of a problem, and the errors were found by running the gate rather than by reading anything.

There was a second-order problem. Several people or processes working in parallel against one shared development database means a full test run measures the database as much as the code. A batch run showed two failures that passed when the same suites were run individually, because another process was writing to the same tenant at the same time.

What we changed

The habit is tedious and it works: run the gates yourself, at the end, against the tree you are about to hand over. Lint, type check, migrations forward and back, the test suites, the browser suite. The value is not catching anyone out. The report and the gate disagree often enough to be worth twenty minutes, and each disagreement is otherwise found by somebody downstream at a worse moment.

A full-suite result now carries its conditions. An uncontended run is a claim about the code, and a contended one is a claim about the afternoon. The reports that were worth reading all had the same contents:

  • The command and its real output: suite names and counts, not the word pass.
  • What was not verified. Several rounds of mobile work ended with an explicit statement that no screen had been driven on a device.
  • Deviations from what was asked, including files touched outside the stated scope.
  • Adjacent defects found and not fixed, logged precisely enough for someone to pick up.

What it did not fix

More and more work is done by processes that report on themselves. A generated summary saying everything passed is written with the same confidence whether or not it did, and it is fluent either way. A badly written report from a tired person carries signal in its raggedness; a generated one does not.

Re-running the gates does not remove that risk. It moves it to a check somebody has to remember to do. I do not exempt myself from this, and the first paragraph is the reason.

What to ask your own team or supplier

  • When a status report says tests pass, which command produced that, and was it run against the code we are accepting?
  • Does the report say what was not verified, for example no screen driven on a real device?
  • Was the test run made while other work was touching the same environment?
  • Who re-runs lint, type checks and migrations at the end, and is that written down?
  • Where a pipeline or wrapper script sits between the tool and the result, which exit status is being read?

Where this ends up

A report is a claim about work and a gate is a measurement of it. That is the rule behind AI-first delivery: verification is a separate act from production, the evidence is the running system rather than the description of what was done, and a check that can fail is worth more than a report saying it passed.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch