AI in your quoting system without sending prices out
An AI helper for a quoting system needs to read every rate and margin you have. So where it runs matters more than how clever it is. A design, not a shipped feature.
You want an AI helper inside your quoting system, something your estimators can ask “why is this price coming out low?” The question you should ask first is where your pricing goes when it answers. To be useful, the helper has to know every pricing rule, every material rate, every labour charge, and the gap between your cost and your selling price on each.
For a fabrication business, that is not data about the business. It is the business. It is what a competitor would need to underbid you on any job, and it took decades to build up. If the helper is a service run by someone else, all of it travels to their servers, and you cannot take it back.
So the first design rule was not accuracy or speed. It was that this knowledge never leaves the machine it lives on. That rules out every hosted AI service, and it makes everything after it harder.
What was actually going on
The system runs on a modest shared server: a couple of processor cores, no graphics hardware for AI, single-digit gigabytes of memory and other applications on the same machine. A language model on that is slow. An answer of any length takes tens of seconds, and the request has a two-minute limit before it gives up. There is no room for a large model or a long conversation.
That is the trade, stated plainly. A hosted service would be faster, cheaper to run and better at answers. It is ruled out by what the helper would be reading, and the right response is to design around the limit rather than quietly relax it.
The design that makes it viable treats the language model as the last resort. Most questions an estimator asks are lookups: what does this rule do, which products use it, which rate does this code point to. Each of those is an exact query against the real records, instant and needing no model. Only the open-ended questions, such as how to set up a new product, need language. So the helper looks things up first, then searches the documentation and earlier approved answers, and uses similarity matching only as an optional third layer. The model writes from what those return, when it runs at all.
The most important sentence in the design is short: if the local AI service is unavailable, the helper says so and stops. It does not fall back to a hosted service “just for availability”, because a silent fallback is exactly how the pricing would leave the machine, on a day nobody is watching, caused by an unrelated outage.
Two boundaries also had to be drawn. Everything the helper knows is scoped to the company it belongs to, filtered when the answer is looked up, not merely asked of the model. And the helper is read-only. Approved question-and-answer pairs can improve its own notes, but prices, rates and templates change only through the existing admin screens, with their permission checks. There is no route by which a conversation changes a price.
What we recommended
Nothing has been built or changed in the product. This is a specification, not a shipped feature. It was written against a verified picture of the server it would run on, with model sizes, storage locations and failure behaviour recorded, and then stopped there.
What the specification does settle is the order of decisions: what the helper may read, then where it runs, then what happens when it cannot run, all written down before anyone implements anything.
What is still open
The honest unknown is speed. I am reasonably confident about the structure, meaning the look-up-first layering, the stop-rather-than-fall-back rule and the two boundaries. I am much less confident that the answers would arrive fast enough on that hardware.
My expectation is that the lookups and the documentation search would carry most of the value, and the written answers would be slow enough that people route around them. If so, the right move is to ship the lookups and drop the generation, not to relax the rule that made generation slow.
The pattern, for anyone adding AI to a system that holds your margins
Decide what the helper will read before you decide where it runs, because the first answer fixes the second. If it reads your commercial model, it runs on your own machine or not at all.
Then ask what happens when it breaks. A safe design degrades into having no helper, never into the same helper with the protection removed. And ask whether it can only read, or also write: a helper that can write inherits every permission question the rest of your software answers, except that it now depends on interpreting English.
Where this ends up
The pricing model in question is the one Sazinga Quote holds: material rates, labour charges and the formulas that turn dimensions into a price. The assistant described here is not built. The rule about where that data may be handled is the part that already applies.