Letting staff ask the system in plain English, safely
A plain-English assistant over a business system must not show a junior the finance figures. We built the limits into what it can see, not what it is told.
You want your staff to be able to type a sentence into your business system and get an answer. Which sites are free next month? Log this expense against that location. What does this client owe? What you do not want is a junior being able to see the finance figures, or changing something by accident because the assistant misunderstood a sentence.
That was the requirement for the operations system we built for Gold Sign Media, an outdoor-media business. It is not a chatbot. It is a way of driving the system’s own operations by typing. The risk is easy to state: the system has rules about who may see or change what, and now a language model sits in front of it. If it can be talked into ignoring a rule, a rule that protects your figures is only as strong as its politeness.
What was actually going on
The common answer, and the wrong one, is to put the rules in the instructions given to the model: “do not reveal financial information to users who are not finance staff.” Instructions are guidance, not enforcement. A model can be talked out of them by someone who is trying, and occasionally by someone who is not. Treating them as the security boundary is a category mistake.
The rules therefore had to live in the structure. We built it in four layers, and only the first three carry any weight.
What we changed
The first layer is the most important. The list of actions handed to the assistant is built fresh for each request from the asker’s own permissions. A user who may not view invoices is never given the invoice actions. The assistant cannot be persuaded to use something it cannot see, and if it names one that is not in its list, that is simply an error. Attacks that work by changing what the model decides to do have nothing to work on when the forbidden action is not among the choices.
The second layer is that each action runs the system’s real code, in the same way a screen would, not a lightweight copy written for the assistant. The permission check built into the real code runs again, so even if the first layer were wrong, the system itself would still refuse. It also means there is one implementation of every operation. A copy written for the assistant would be a second implementation, and second implementations are where permission gaps live.
The third layer is that nothing which changes data happens straight away. An action that would change something returns a proposal: what would be done, in plain words. The person confirms it through an ordinary, separately permission-checked step. A misread sentence costs a confirmation question, not a corrupted record. A wrong answer to a question wastes a moment, while a wrong change has to be found and undone, and might not be found.
The fourth layer is the instructions: the assistant acts only through its actions, does not answer general questions, does not invent data, and is told the currency and today’s date, which it would otherwise guess. That layer improves behaviour and enforces nothing.
To turn a phrase such as “the site on the main road” into a real record, the assistant uses a search over the existing records and then uses the code it gets back. The records are always current, and the search answers exactly, where a separate index of them would need building and keeping in step with data that changes daily.
One fault turned up that is a good warning for anyone building this. The assistant had been reporting the size of one page of results as the total number of records. It gave a count that was wrong by a large factor, confidently, because nothing it received said more existed. Making the answers shorter to save cost fixed it as a side effect: every list now says how many there are in total and how many it is showing.
What it did not fix
The assistant can only be as right as the information it is handed, and it will report what it can see as if it were everything. The permission rules it obeys are the system’s own, so a mistake in those rules is a mistake the assistant repeats. The structure limits what a wrong interpretation can do. It does not make every interpretation right.
The pattern, for anyone letting staff talk to their systems in plain English
Ask one question of whoever builds it: what would have to be true for a user to do something they are not allowed to do? If the answer includes “the assistant would have to ignore an instruction”, the design is not finished.
A sound answer sounds like this: the action would have to be in a list built from that user’s own permissions, and the system behind it would have to skip its own check, and the user would have to confirm a change described to them in plain words. Then the assistant is not the last line of defence, which is where it should not be.
Where this ends up
The assistant described here runs over Sazinga AdBoard, so each person sees what they may see, and a change only happens once a person has confirmed it in words they can read.