A chatbot over your company data can be confidently wrong
Before building an AI assistant over an orders database, we found 2,667 approved orders it would have reported as cancelled, with no error.
You want to ask your business system questions in plain English. How many orders did each representative close last month? Which dealers have nothing approved? Somebody connects an AI model to the database and the answers arrive instantly, in a confident tone.
An assistant that picked the wrong status table would have told an owner that 2,667 approved orders were cancelled. No error, no warning, and a plausible-looking figure. A decision made on that answer would have been made on the wrong number, and nothing on the screen would have hinted at it.
What was actually going on
An AI model can read the names of the tables and columns in your database. What it cannot read is the knowledge that only ever lived in the code and in the heads of the people who built it: which column in one table refers to which person in another, when nothing in the database says so.
So before building anything, we catalogued what the existing system really does. We went through every query the service runs and recorded how its tables are actually connected. Of 45 distinct connections, 24 are backed by a rule in the database itself. 21 are not. Those 21 run in production every day, and they are invisible to anything that reads the database alone, including an AI model and every new engineer for their first fortnight.
A model that does not know them invents one, and an invented connection usually runs. It returns rows. They are the wrong rows, in a plausible quantity, and nothing errors.
The status problem was the same kind of trap. The system has two tables of order statuses. One is live and is what orders point at. The other is an old leftover that nothing points at, but its numbers collide with the live one: the number 2 means Approved in the live table and Cancelled in the old one. There are 2,667 orders sitting on 2. Choose the wrong table and every one of them is reported as cancelled.
Two more traps turned up. There are two different things called region, one geographic and one a business grouping, so a bare question about “region” is genuinely ambiguous and the assistant should ask. And a list of job titles written into the code disagrees with the live database, every number shifted by one. The database wins, and the code had been wrong all along, silently, for whatever it is used for elsewhere.
What we changed
Each connection in the catalogue carries the kind of evidence behind it: a real rule in the database, a connection the code actually uses, or merely two tables mentioned together. That last kind is recorded but labelled as the weakest, and is not offered as a route, because presenting it as one would teach the model to invent exactly what we were trying to stop.
Every table name was then checked against the live database. That caught names the extraction had invented itself, such as a plural form that does not exist. A wrong table name is worse than a missing one: missing produces a question, wrong produces an answer. Five names are still reported as unresolvable, and listed, so the catalogue shows where it stops.
The values that mean something, such as order statuses, roles and regions, are read live from the database and given to the model, so it never has to guess what a 2 means.
What it did not fix
This is a generated catalogue and a prototype, and it has known edges. The hints about which tables each function touches are broad and over-report, so they are a starting point, not a contract. Connections built while the system is running cannot be found this way. And it describes what the code does, not what it ought to do.
It also turned up something nobody asked for. Of 266 read-only endpoints, 66 carry an explicit permission check and 200 do not. All require a login, but the decision about who may see what lives inside each handler, where it has to be read case by case. That was reported as a finding rather than quietly corrected.
The pattern, for anyone thinking of putting an AI assistant on their data
Ask what the assistant has been told about how your records connect, and who checked. If the answer is “it reads the database”, it will guess wherever your system relies on habit. Ask, too, what it does when a question is ambiguous. The right answer is that it asks.
Most systems of any age carry knowledge that exists only in the code and in people’s heads. Writing it down and checking it against the database is worth doing even if you never build the assistant. We only did it because a model needed it, which is a poor reason to have waited.
Where this ends up
This catalogue is the first step of an engagement rather than a by-product of one, and it is delivered as a report you keep whatever you decide afterwards. See AI data assistants.