Was your supplier testing on your live data?
End-to-end tests could only run on the live system, so a defect was found by demonstrating it on real quotations. One suite now gates the push and checks the deploy.
For a while the only place the end-to-end tests could run was production. There was no local environment with realistic data, so a change went out and then a browser-driving suite confirmed it against the live system, on the customer’s real records.
That works until the day something is wrong. Then you have found a defect by demonstrating it on live data, and the customer’s quotations are the test fixtures. If your supplier tests on your live system, every defect they find, they find on your records.
What was actually going on
The suite already drove a real browser against a real deployment. It only assumed which one. The fix was smaller than it sounds: the suite reads the interface address and the API address from the environment. In production they are the same address. Locally they are not, so they are supplied separately, with the API address defaulting to the interface address.
That leaves exactly one suite, not a pre-push suite and a post-deploy smoke test. Two suites reliably check different subsets of behaviour and drift apart over a year until neither is trusted. Here the tests that gate the push are the tests that verify the deploy.
The local stack is a restored snapshot of the production database in a separate local database, brought up to the current schema. Real records rather than seeded fixtures caught something real. A change to how dimensions are converted between measurement systems had a write path I had missed, one editing surface that stored a value without converting it. Against fixtures that produces a number, and a number is a number. Against real records it produced a work table at a twenty-fifth of its correct price, sitting next to forty other lines that looked right.
What we changed
The snapshot carries obligations. It lives locally, it is not shared, and anything done against it is done knowing it contains a real business’s commercial data.
The document renderer is a native dependency installed on the server and not on every development machine, so a local API without it cannot produce the printed quotation and returns a specific rejection. The suite tolerates exactly that endpoint and that client-error response. It does not tolerate a server error from the same endpoint, which would mean the export path is broken, and the document template is exercised separately by rendering it directly.
Because the same suite runs against production, it creates real records there. Every quotation it creates carries a recognisable prefix and is deleted on teardown. The prefix matters more than the deletion: when teardown fails, leftovers must be identifiable by someone who was not there, because a test record that looks like real data ends up in a revenue report.
A deploy follows the same routine each time: back up the database, deploy the code, apply migrations, then verify three things. The health endpoint responds, the schema version matches, and one specific piece of data the change was meant to affect is checked directly. The third check is what makes it a verification, because a service that started and a migration that ran prove nothing about intent.
What it did not fix
The pre-push run happens against a snapshot taken at some point in the past, so a change that interacts with data created since can behave differently. There is no continuous integration running any of this automatically. It is a discipline executed by a person, and disciplines executed by people are skipped under time pressure.
It removes the class of failure where a change is first exercised on live data. It does not remove the class where nobody ran the tests.
What to ask your own team or supplier
- Where do the end-to-end tests run, and have they ever run against our live system?
- If they run there, how are test records labelled and removed, and how would a leftover be found?
- Is the pre-release suite the same suite that checks the deploy, or two that can drift?
- Is a copy of our production data held locally, who can reach it, and is that covered in the contract?
- What does the deploy check prove about this change specifically, beyond that the service started?
Where this ends up
The same reasoning sits under AI-first delivery: prove the check before trusting it, because a test you have not seen fail is not yet a test, and read the state back rather than accepting a success message as evidence that it changed.