AI & Machine Learning
Hire Data Engineers
Engineers who build the pipelines and schemas everything else depends on, and keep them trustworthy.
Data engineering earns its place at the point where reporting stops being a query somebody runs and starts being something the business decides on. Sazinga AdBoard, Rentals and Comply all sit on PostgreSQL, and the schemas, migrations and extract jobs behind them are maintained by the same engineers we hire out, which is why that work is treated as product code rather than as scripts.
A data engineer’s real product is not a pipeline, it is a guarantee: that a number can be reproduced, traced back to its source, and shown to be wrong when it is wrong. Everything else — the orchestration tool, the warehouse, the modelling layer — is implementation detail chosen in service of that.
The failure mode that never raises an error
The expensive failures in data work do not throw. They succeed with the wrong result, or they stop happening at all, and nobody notices because nothing was ever going to shout.
Our own contact pipeline lost two weeks of enquiries that way. Submissions were being stored
correctly the whole time; the notification step was not running, and a stored enquiry that had
never been emailed looked exactly like one that had. A hundred and fifty-three of them accumulated
before anyone realised. The fix was not more logging. It was a notified_at column and a view that
selects rows where it is null — an invisible failure becomes a queryable one the moment the
absent step has a column of its own.
That principle generalises to every pipeline we build. Record the fact that a stage completed, not only the data it produced, and make “should have run and did not” a row you can count rather than a silence you have to notice.
Four schema decisions that are expensive to reverse
Most of a schema can be changed later. A short list cannot, and it is worth slowing down for.
- Keys that have escaped. One of our tables keeps a timestamp-plus-hex identifier rather than a serial, because that string is already printed as “Ref:” in every notification and is the filename of every stored attachment. A tidier key would have orphaned both. A key is only yours to change while it lives inside the database.
- Type strictness at the ingest edge. A visitor’s IP is stored as
textrather thaninet, deliberately: a malformed value must never be the reason a lead fails to store. Choose the type by what should happen when the data is wrong, not by what the data ought to be — permissive at the boundary, strict in the models downstream. - How money and rates are recorded. Store the inputs and the rule alongside the answer, and keep the precision of a rate separate from the precision of an amount.
- Tenancy and time. Which column separates one customer’s data from another’s, and whether a timestamp means the event or the load, are both very hard to retrofit once reports depend on them.
Reruns, backfills and the things only testing makes true
A pipeline is finished when running it twice produces the same result as running it once. That is
the whole of idempotency, and it is why our own reconciliation job moves each processed file into a
migrated/ directory rather than deleting it — you cannot rerun a job against evidence you threw
away.
Backfills deserve more suspicion than they get, because a default applied to historic rows is a decision, not a formality. Migrating our stored submissions into Postgres, the historic spam was rescored on the way in and written as held rather than defaulting to clean, which is the difference between a migration that preserves the truth and one that quietly manufactures it.
And a backup is a claim until it has been restored. Ours was verified by restoring into a scratch database and counting rows, not by confirming a file existed on disk with a plausible size.
Where hiring a data engineer is the wrong move
If nobody in the organisation agrees what “active customer” means, a pipeline will encode one person’s answer and make the disagreement permanent. Settle the definitions first; that work is analysis, not engineering.
If the data fits comfortably in the operational database and the actual complaint is that reports are slow, the honest answer is often three indexes and a materialised view, delivered by the application team in a fortnight. A warehouse programme started to solve that is an expensive way to avoid reading a query plan.
And if there is no owner for the alerts, do not build the alerting. Data quality checks that fire into a channel nobody reads are worse than none, because they create the impression of supervision.
What we ask in an interview
How does your pipeline behave on a rerun, and what did you do the first time it did not. What happens when an upstream file arrives late, arrives twice, or arrives with a column removed. How did you find out the last time a number in a report was wrong — from a check, or from a person.
Engineers who have carried a pipeline in production answer with specifics about idempotency, reconciliation and watermarking. The rest describe the happy path, in detail, and at length.
Teams are built for companies in the United States and the Gulf — the UAE, Saudi Arabia, Qatar, Kuwait, Bahrain and Oman. The engineers are in Pune, which matters mostly for the clock. Dubai is ninety minutes behind us and Riyadh two and a half hours, so a Gulf team shares almost the whole working day. New York is nine and a half hours behind, so American engagements run on a written handover and one fixed overlap window rather than on a standing call — a real constraint, and better stated than discovered.
What these engineers do
- PostgreSQL schema design, indexing, partitioning and query plans under real write load
- Batch and incremental pipelines in Airflow or Dagster with idempotent reruns
- dbt models, tests and documented lineage in place of undocumented SQL sprawl
- Change data capture and warehouse loading into BigQuery, Redshift or Snowflake
- Data quality checks that block a bad load rather than reporting it the next morning
Delivered AI-first
AI assistance is used for the parts of data work that are structurally repetitive - generating dbt model and test stubs from a source schema, drafting pipeline DAGs, translating stored procedures into versioned SQL, and writing the migration scripts that accompany a schema change. Correctness is established by tests and by reconciliation against source counts, never by the assistant asserting it. Review of schema changes and migrations is unchanged. The measurable effect is throughput per engineer, not fewer reviews.
Built with Data Engineering at Sazinga
These are our own production applications, not client references — which is why the engineers have operated them, not just written them.
Other AI & Machine Learning roles
AI/ML
Machine learning engineers who put models into production and keep them working, not notebooks that never ship.
LLM & RAG
Engineers who build retrieval-augmented systems that answer from your own data and admit when they cannot.
MLOps
Engineers who make model delivery boring — versioned, reproducible, monitored and reversible.
What an unfilled engineering role costs while you hire — worked out on your own numbers.
Tell us what the Data Engineering work is.
Roughly what it involves, the seniority you need, and when it has to start. We will say what it takes to staff it, or say honestly that we are not the right people for it.
Thank you — that has reached us
Your enquiry is with the team. We read every one ourselves and normally reply within one working day.
If it is quicker to talk, reach us directly:
While you wait — the platform overview covers what each application does, and Insights is our writing on building this kind of software.