Let's talk
data-modelling

The till would not bill, because another business was busy

A till refused to bill for no reason its owner could see. The invoice number came from a clock, repeated every 16 minutes and was shared with every other customer.

·

A tablet on a stand at a till showing the Add Sale screen in Sazinga Factory.

The till refused to bill. Not every time — most sales went through — but now and then a sale at the counter would fail, the operator would try again, and eventually it would go through or it would not. Nothing about the sale was unusual. The stock was there, the customer was there, the price was right. The software simply would not issue the invoice.

If you run a shop or a plant with several tills you have probably seen a version of this and put it down to the network, or the server, or the time of day. Here it was none of those. The till was failing because a different business, using the same software from somewhere else entirely, happened to be busy at that moment.

What that costs is a customer standing at a counter while the operator retries; a support ticket from a business that has done nothing out of the ordinary; and, for whoever runs the software, a fault that cannot be reproduced because it depends on what somebody else was doing at the time. It also cost every other till in the shop up to half a second of waiting behind the stuck one, for a reason that comes later.

What was actually going on

Every sale needs an invoice number: the number a customer reads out over the phone and an accountant types into a return, so it has to be unique. This system made it by looking at the clock. It took the current time in milliseconds, kept the last six digits, and put the year in front.

Six digits gives a million possible values, and the clock moves through one of them every millisecond. So the whole sequence repeats every million milliseconds, which is sixteen minutes and forty seconds. The number issued right now will be issued again a quarter of an hour later, and again a quarter of an hour after that, all year. The year in front does not help, because every sale that year shares it.

The only thing standing between the system and a constant stream of duplicates was that two tills rarely bill in the same millisecond. Rarely. Across every till in every business using the system, for twelve months.

Here is the part that turns a rare clash into another business’s problem. The rule that says “no two invoices may share a number” was written across the whole system, not per business. Every other record in the system is kept apart by company; the invoice number was the one value shared across the whole estate. A busy customer, generating lots of numbers, raised the chance of a clash for everybody, and the business that experienced the failure was whichever one happened to be second.

The author knew. Above the numbering code sat a note saying the clock-based scheme avoided clashes. Below it sat a loop that caught the clash when it happened, made up a different number, waited fifty milliseconds and tried again, up to ten times. Both of those cannot be true. If the scheme avoided clashes there would be nothing to catch. The retry loop is the author’s own evidence that the note is wrong, and it had sat there, unremarked, for as long as the module existed.

The retry was worse than the first attempt. On a clash it rebuilt the number from the last three digits of the clock plus three random digits. Three digits of clock repeats every second, so the fallback was a thousand-way random draw against whatever was already in the table.

And the fifty-millisecond wait happened while the sale was half-written, with the system holding the stock records the sale was about to reduce. Fifty milliseconds is nothing to a person and a long time to a lock. Ten attempts is half a second during which every other till trying to sell the same product waits behind the stuck one. That delay only ever arrived under load, which is exactly the moment you least want to add delay.

The same mistake was in the purchase numbers, for the same reason: the rule was written when the first module was built, and nobody revisited it when the system started serving more than one company.

What we changed

Invoice numbers now come from a counter: one per company, per document series, per year, moved on and handed out in a single step that the sale itself owns. If the sale fails, the number goes back. It gives numbers a person can read, a sequence that means something, and it only ever competes with other sales in the same company and series, which is the only competition the business rule ever implied.

The uniqueness rule now says what was meant: no two invoices within one company may share a number. Another company’s busy afternoon can no longer touch your till. The retry loop went away rather than being tuned.

What it did not fix

Replacing the generator is easy and migrating the numbers already issued is not. Those numbers are printed on invoices customers hold. The change fixed the rule and the generator for new documents, left the history untouched, and accepted that the series has a visible break at the date of the change — explicable, and much better than a series that quietly repeats every sixteen minutes.

The mechanism, for anyone checking their own system

The clock scheme was chosen to avoid a counter, because a counter needs coordination and coordination looked slow. That is a reasonable instinct and the wrong trade here, because a document number is not a high-throughput identifier. It is issued once per sale, by a person, at a till. A counter row per tenant, per series, per year, incremented and returned in one statement, is atomic and transactional, and contends only with sales in the same tenant and series.

If a readable sequence genuinely is not required, the honest alternative is the opposite extreme: a random identifier wide enough that clashes are not a design consideration at all. What does not work is the middle — a short value derived from a clock, patched with a retry, and treated as if the retry were an optimisation rather than an admission.

Two rules follow. A uniqueness constraint is a statement about the scope in which a value must be unique; if your product serves more than one company, a constraint that omits the company column is asserting something you do not mean, and every such rule should be read aloud with the company missing to see whether the sentence is still one anybody asked for. And never sleep inside an open transaction: if a retry needs a pause, the unit of work has to end first and be repeated whole. A retry that keeps the transaction open is not a backoff, it is a hold.

The habit worth taking is to read every retry loop as a confession. A loop that catches a duplicate is telling you the author knew the generator could produce one, and the retry count is a rough estimate of how often they expected it to.

Where this ends up

Document numbering in Sazinga Factory is drawn from a counter the transaction owns rather than from the clock, which is what lets the retry loop go away instead of being tuned.

This came out of building Sazinga Factory

From raw material to finished batch, with the yield accounted for. The problem above is one we met while building it, and what we did about it is in the product.

If you run something like this, there is one thing you can do without a call: send one batch record.