Let's talk

AI & Machine Learning

Hire AI/ML Engineers

Machine learning engineers who put models into production and keep them working, not notebooks that never ship.

Machine learning earns its place when a decision is made too often for people to make it well every time — demand forecasts, anomaly flags on operational data nobody has time to read, ranking work that would otherwise be a spreadsheet. It stops earning its place the moment the model cannot be retrained, monitored or explained to the person whose job it changes.

The most useful model on a given problem is frequently not a trained one. One of our own production classifiers labels vehicle photographs into two dozen categories — odometer, fuel gauge, signature sheet, damage, each side of the car — and it has no training run at all. A general vision model embeds a sample of the photographs a human has already labelled, the vectors for each category are averaged into a prototype, and a new photograph is scored by how close it sits to each prototype. No fitting, no learning rate, no overfitting to worry about, and a new category is added by supplying examples.

Abstention is a product decision, and it is the one that makes it usable

The score that matters is not the top similarity; it is the gap between the best match and the second best. A small gap means the model cannot tell two categories apart, and the correct behaviour is to say nothing and leave the item for a person.

Measured on our own held-out photographs, that single threshold is the whole product:

  • Auto-apply everything: 100% coverage at 70.6% accuracy — worse than useless, because every wrong label now has to be found.
  • A modest margin: 55% coverage at 86.9%.
  • A strict margin: 30% coverage at 96.1% accuracy, with the remaining 70% routed to a human.

The last row is the one that shipped. A classifier that answers less often and is right when it answers is worth more than one that always answers, because the cost of a wrong label here is someone finding it later in a dispute, and the cost of no label is a person spending four seconds.

The per-category results say the same thing in more detail. Odometer, fuel gauge, signature and checklist images are essentially solved. Damage photographs sit around a third correct, and the left and right sides of a car are close to indistinguishable to a general vision model — they are mirror images. That produces a small margin, which routes them to a person, which is exactly the behaviour we want from a failure we cannot currently fix.

Evaluation that would embarrass you if it were wrong

The held-out set is taken from an offset past the samples used to build the prototypes, so being disjoint is a property of the code rather than a claim in a document. The evaluation prints per-category accuracy, the abstention rate and the categories most often confused with each other, because the confusion pairs are what tell you whether the next improvement is more data, a bigger model, or a change to the category definitions.

The honest baseline question comes first, though. Before any of this, someone should ask what accuracy a trivial rule achieves — most common class, a keyword, a timestamp ordering. A model that does not clearly beat that is a maintenance liability with a research paper attached.

Running inference next to something that people depend on

The production machine here also runs the live API and two databases, and has no GPU. That constraint produced the design and it is the part we would repeat anywhere.

Inference runs off the production box entirely. It reads what it needs through read-only queries with the session forced into read-only mode, pulls the images from object storage rather than through the application, and watches the server’s load average — pausing when the machine is busy. Writes are throttled to a couple per second and touch only records a human has not already labelled, so the automated pass can never overwrite a person’s judgement.

And it runs as a dry run by default, producing a file of proposed labels with their scores. Applying them requires a flag. The default behaviour of a model that writes to a production database should be to write nothing.

Where machine learning is the wrong tool

Where the rule is knowable. If a domain expert can write the decision down in ten lines, write the ten lines: they are testable, explainable and free to run.

Where the labels do not exist and nobody is willing to create them, because the labelling effort is the project whether or not anyone plans for it. Where the decision is high stakes and rare, so there is no volume to learn from and no tolerance for a wrong answer. And where nobody will own the model after launch — an unowned model does not stay still, it decays.

What we interview for

Look past framework familiarity. Ask how a candidate designed a holdout set, what baseline they beat and by how much, and what happened to the model six months after launch. Engineers who have owned a model in production talk about drift, feedback loops and rollback. Those who have not talk mostly about architectures.

Then the one that reveals judgement: tell me about a case where you decided not to use a model.

Teams are built for companies in the United States and the Gulf — the UAE, Saudi Arabia, Qatar, Kuwait, Bahrain and Oman. The engineers are in Pune, which matters mostly for the clock. Dubai is ninety minutes behind us and Riyadh two and a half hours, so a Gulf team shares almost the whole working day. New York is nine and a half hours behind, so American engagements run on a written handover and one fixed overlap window rather than on a standing call — a real constraint, and better stated than discovered.

What these engineers do

  • PyTorch 2.x and scikit-learn for training, evaluation and error analysis
  • Feature pipelines, drift detection and retraining triggers that actually fire
  • Model serving behind FastAPI, ONNX Runtime or Triton within real latency budgets
  • Forecasting, classification and anomaly detection on messy operational data
  • Honest evaluation - holdout design, leakage checks and baselines worth beating

Delivered AI-first

AI assistance sits in the delivery loop rather than in the model itself. Engineers use it to scaffold experiment harnesses, dataset loaders, evaluation code and the boilerplate around serving and monitoring, which leaves more of their time for problem framing and error analysis. Every training script, metric definition and inference path is still read line by line in review. The measurable effect is throughput per engineer, not fewer reviews.

Tell us what the AI/ML work is.

Roughly what it involves, the seniority you need, and when it has to start. We will say what it takes to staff it, or say honestly that we are not the right people for it.

A person reads every enquiry and replies within one working day.