AI & Machine Learning
Hire LLM & RAG Engineers
Engineers who build retrieval-augmented systems that answer from your own data and admit when they cannot.
Retrieval-augmented generation earns its place where the answer already exists somewhere in your documents, tickets or records, but finding it costs a person twenty minutes. It fails, expensively and publicly, when it is bolted onto a corpus nobody has curated, with no evaluation set and no honest measure of how often it is wrong.
Retrieval failure and generation failure are different bugs with different fixes, and a system that cannot tell them apart cannot be improved. If the passage containing the answer was never retrieved, no amount of prompt work will help. If it was retrieved and the answer still came out wrong, the index is not the problem. The first diagnostic any of these systems needs is one that reports, for a question with a known answer, whether the supporting passage made it into the context at all.
Structured output constrains shape, not sense
Every serious provider will now accept a schema and return JSON that matches it. That is genuinely useful and it is routinely mistaken for validation.
A generated quiz question in one of our systems is schema-valid with a prompt that is not a string we can render, with fewer than two distinct options, or with an answer index that points at the wrong option after blank entries are removed. All three pass the schema. So everything is re-validated on our side: types checked strictly, boolean values rejected where an integer index is required because booleans are a subtype of integer in the language, and any item whose answer index shifted during cleaning is discarded rather than kept — silently keeping it would mark a wrong answer correct in front of a child.
Two related rules. Invalid items are skipped rather than fatal, because one bad entry used to fail a whole batch. And the JSON extractor is deliberately forgiving: strip a wrapping code fence, take from the first brace to the last, and only then parse.
Truncation is the failure nobody designs for
Output caps produce cut-off responses, and a cut-off response is not an error — it arrives with a success status and looks like an answer. In a system where the model wraps its draft in delimiters, the unterminated case has to be handled explicitly, or the raw marker reaches the user’s screen.
Two mitigations that cost nothing: tell the model the length it is aiming for in the prompt, so it composes something that finishes, and return a truncation flag alongside the text so the caller can decide rather than guess.
There is a subtler version. On one provider, reasoning tokens counted against the same output cap, so a capped request spent nearly its entire budget thinking and returned about fifteen characters. Setting the cap without disabling the reasoning budget looked like a broken model.
Failover between providers, without failing the user
Multi-provider support is easy to describe and full of traps. What we run to is a small set of rules.
Retry once, and only for network and timeout errors. A 4xx or a 5xx is a real answer from the provider and retrying it just costs latency. Advance to the next model in a fallback list only on a genuine “no such model” signal, never on a rate limit or an authentication failure, and keep the resolved model cached separately for text and for vision so a successful text call cannot promote a model that cannot see images.
Treat three provider failures as distinct types, because the right response differs: an invalid key, a rate limit carrying a retry interval, and a billing failure. Putting a billing-dead key on a cooldown rather than marking it makes the situation worse, not better.
And keep the client timeout under the reverse proxy’s read timeout, so the edge never gives up before the application does. A 504 from a proxy carries none of the information a handled timeout does.
Cost is a design parameter, not an invoice
Budgets are enforced per feature, not discovered monthly. A daily token cap measured on a fixed UTC boundary rather than a local one, so the same person cannot spend the allowance twice from two contexts. A nominal charge applied to any call whose provider reports no usage figures, so an unmeasured runaway loop cannot be free. And an explicit maximum output length on every request — one routed model defaulted to its own very large limit and produced a billing error that looked exactly like a bad key.
Where retrieval augmentation is the wrong choice
Where the corpus is small enough to fit in the context window. Retrieval adds an index to maintain and a failure mode to debug for no benefit; put the documents in the prompt.
Where the question is a database query in disguise — “how many open tickets does this customer have” is SQL, and answering it by embedding ticket text is a worse database with a language model attached.
Where nobody will curate the corpus. Retrieval faithfully surfaces the outdated policy document that nobody deleted. And where the domain cannot tolerate a wrong answer and there is no route for the system to say it does not know.
What we interview for
Ask what their golden set looked like, how they told retrieval failures apart from generation failures, and what the system does when confidence is low. Engineers who have run one of these in production will talk about chunking regrets, index refresh and cost per query. Anyone who only talks about model choice has not yet run one.
Then ask what their test suite mocks, and what it cannot. The answer we want to hear is that transport is mocked for speed and a small live check runs on a schedule — because the surprises that matter are all provider behaviour that no mock will ever reproduce.
Teams are built for companies in the United States and the Gulf — the UAE, Saudi Arabia, Qatar, Kuwait, Bahrain and Oman. The engineers are in Pune, which matters mostly for the clock. Dubai is ninety minutes behind us and Riyadh two and a half hours, so a Gulf team shares almost the whole working day. New York is nine and a half hours behind, so American engagements run on a written handover and one fixed overlap window rather than on a standing call — a real constraint, and better stated than discovered.
What these engineers do
- Retrieval pipelines with pgvector, Qdrant or OpenSearch and hybrid keyword reranking
- Chunking, embedding and index refresh strategies for documents that keep changing
- Evaluation harnesses with golden sets, faithfulness scoring and regression gates in CI
- Tool calling, structured output and agent loops designed to fail safely
- Token cost, latency and caching budgets tracked per feature rather than per invoice
Delivered AI-first
This is the one stack where the technology and the delivery method overlap, and the distinction matters. AI assistance is used to generate prompt scaffolding, retrieval test fixtures, evaluation datasets and adapter code between model providers. It does not decide what counts as a correct answer. Golden sets are written by engineers who understand the domain, and every change to a prompt, retriever or ranking function goes through review and an evaluation run before it ships. The measurable effect is throughput per engineer, not fewer reviews.
Other AI & Machine Learning roles
AI/ML
Machine learning engineers who put models into production and keep them working, not notebooks that never ship.
Data Engineering
Engineers who build the pipelines and schemas everything else depends on, and keep them trustworthy.
MLOps
Engineers who make model delivery boring — versioned, reproducible, monitored and reversible.
What an unfilled engineering role costs while you hire — worked out on your own numbers.
Tell us what the LLM & RAG work is.
Roughly what it involves, the seniority you need, and when it has to start. We will say what it takes to staff it, or say honestly that we are not the right people for it.
Thank you — that has reached us
Your enquiry is with the team. We read every one ourselves and normally reply within one working day.
If it is quicker to talk, reach us directly:
While you wait — the platform overview covers what each application does, and Insights is our writing on building this kind of software.