Let's talk
ai

362 pricing rules to check, and nobody gets past forty

A fabricator's 362 spreadsheet formulas each needed a sign-off from somebody who knows the trade before a quote could be trusted. The review stalled at item forty.

·

A laptop on a steel desk in a works at night under one desk lamp; beside it the Sazinga Quote formula check showing a rule, its steps and a sample price.

Your quoting has just come off the spreadsheet. Every formula the business used to price a job — 362 of them — has been carried across into the new system, and before you let it quote a single customer you want somebody who knows the trade to look at each one and say yes. That somebody is you, or the one estimator you trust, and both of you have a fabrication shop to run.

So the checking starts on a Tuesday evening, gets to about item forty, and stops. It is not that the rules are hard to judge. A person who prices stainless work all day can tell in a second whether a number is in the right region. It is that there are 362 of them in no particular order, no way to see how far you have got, and every one you want to correct means leaving the list and coming back to find your place gone.

What that costs is simple to state. Until the review is finished the system cannot be trusted to quote, so the spreadsheet stays in use beside it and the money spent on the move earns nothing. And if the review is declared finished when it is not, the first wrong price goes out on a real quotation, and nobody can say afterwards which of the 362 rules a person actually looked at.

What was actually going on

The rules were the business’s own, imported from its own spreadsheet. But the import had repaired some of them and rewritten all of them in the system’s own notation, so a reviewer reading a rule was reading algebra, not a price. To know whether it was right they had to work it out in their head. Multiply that by 362 and the task has no shape: no order to work in, no progress to see, and nothing that says which ones deserve attention first.

What we changed

Three things, and the first matters most.

Every rule is shown with a sample price beside it: the number the rule produces for a typical item, at typical dimensions, in a common grade. Not a real quote — a plausible one. A fabricator reading an expression has to simulate it. A fabricator reading a price for a 48-inch item in a common grade knows immediately whether it is in the right region, because pricing that item is what they do all day. The reviewer is shown an output in the units they think in, not an input in the units we think in.

Second, an automated pass reads every rule first and flags the ones that look doubtful. The list can then be filtered to the flagged ones. That is the real feature — not “the software checked your formulas” but “start with these forty”. It turns an unordered pile of 362 into a short queue and a long tail, and a short queue is a task somebody finishes.

Third, the editor opens inside the review screen. A reviewer who spots a problem fixes it and tests it there, and the list keeps its place. Without that, every flagged rule meant leaving the review, finding the rule in the editor, fixing it, and coming back to a list that had forgotten where you were. That friction is what stops people at item forty, not the difficulty of judging a formula.

Who signs, and who only sorts

This is the decision the whole thing rests on, and it is worth being blunt about.

The automated verdict lives in its own place: a status, a note explaining it, the sample price it saw, and the time it ran. The human sign-off is a separate mark, set by a separate action. They never merge, and neither is ever derived from the other. A rule the automated pass called fine is not signed off. A rule a person signed off does not become “automatically verified”.

The moment a machine’s opinion can satisfy the same box a person’s signature satisfies, you have a system where nobody can tell, six months later, which of your 362 pricing rules were actually looked at by somebody who knows the business. That question will be asked, on the day a price is wrong, and “the system reviewed them” is not an answer anyone accepts.

So the automated pass has authority over one thing only: the order in which a person looks at rules. It changes no price. It writes nothing to a rule. It cannot approve anything. Being wrong in that role is cheap — a false flag costs one unnecessary look, and a missed problem leaves you no worse off than before, which was no review at all. Give a machine authority over attention, not over outcomes. The failure mode of misplaced attention is wasted time. The failure mode of a wrong outcome in a pricing system is an invoice.

Keeping the two marks apart also gives you the one measurement that decides whether the automation is worth having: how many rules it flagged, against how many a person then agreed with.

What it does not do

The sample price is computed on invented dimensions. For most rules that is a fair test. Some rules change their behaviour with size — a few branch on width — and a single representative instance exercises one branch and says nothing about the others. A rule can look perfectly sensible at the sample size and be wrong at every other size.

And the value of the automated pass depends entirely on how much it flags. Flag a tenth and it is a queue. Flag half and it is the original problem with extra steps. That ratio is a property of the data and the prompt, not something the design guarantees, so it has to be watched rather than assumed.

Neither of those undermines the structure. Both mean it is doing less than the words “AI verification” might suggest, and it is better to say so.

How the sample price is made

The representative dimensions are a fixed table of plausible values for each dimension the rules use — a typical length, breadth and height, typical shelf, bowl and door sizes, sensible counts for things that are counted — with a neutral fallback for anything unrecognised. It is entirely made up, and that is fine, because its only job is to produce a number a human can react to. The rule is evaluated against those values and sensible default materials, and the result sits beside the rule’s steps, in order, with the expressions as written and the material slots it draws on.

The rule for anyone building a review tool: show the expert an output they can judge in their own domain — a price, a drawing, a result — rather than the artefact that produced it, and put the means to fix it on the same screen. Review tools fail on friction far more often than on judgement.

Where this ends up

The 362 rules are the pricing formulas Sazinga Quote runs on, out of the spreadsheet they were kept in. The reason each one has to be signed by somebody who knows the trade is that every price quoted afterwards is theirs, not the machine’s.

This came out of building Sazinga Quote

The pricing formula, out of the spreadsheet and under control. The problem above is one we met while building it, and what we did about it is in the product.

If you run something like this, there is one thing you can do without a call: send one quotation you have already sent a customer.