Let's talk

AI & Machine Learning

Hire Data Engineers

Engineers who build the pipelines and schemas everything else depends on, and keep them trustworthy.

Data engineering earns its place at the point where reporting stops being a query somebody runs and starts being something the business decides on. Sazinga AdBoard, Rentals and Comply all sit on PostgreSQL, and the schemas, migrations and extract jobs behind them are maintained by the same engineers we hire out, which is why that work is treated as product code rather than as scripts.

A data engineer’s real product is not a pipeline, it is a guarantee: that a number can be reproduced, traced back to its source, and shown to be wrong when it is wrong. Everything else — the orchestration tool, the warehouse, the modelling layer — is implementation detail chosen in service of that.

The failure mode that never raises an error

The expensive failures in data work do not throw. They succeed with the wrong result, or they stop happening at all, and nobody notices because nothing was ever going to shout.

Our own contact pipeline lost two weeks of enquiries that way. Submissions were being stored correctly the whole time; the notification step was not running, and a stored enquiry that had never been emailed looked exactly like one that had. A hundred and fifty-three of them accumulated before anyone realised. The fix was not more logging. It was a notified_at column and a view that selects rows where it is null — an invisible failure becomes a queryable one the moment the absent step has a column of its own.

That principle generalises to every pipeline we build. Record the fact that a stage completed, not only the data it produced, and make “should have run and did not” a row you can count rather than a silence you have to notice.

Four schema decisions that are expensive to reverse

Most of a schema can be changed later. A short list cannot, and it is worth slowing down for.

  • Keys that have escaped. One of our tables keeps a timestamp-plus-hex identifier rather than a serial, because that string is already printed as “Ref:” in every notification and is the filename of every stored attachment. A tidier key would have orphaned both. A key is only yours to change while it lives inside the database.
  • Type strictness at the ingest edge. A visitor’s IP is stored as text rather than inet, deliberately: a malformed value must never be the reason a lead fails to store. Choose the type by what should happen when the data is wrong, not by what the data ought to be — permissive at the boundary, strict in the models downstream.
  • How money and rates are recorded. Store the inputs and the rule alongside the answer, and keep the precision of a rate separate from the precision of an amount.
  • Tenancy and time. Which column separates one customer’s data from another’s, and whether a timestamp means the event or the load, are both very hard to retrofit once reports depend on them.

Reruns, backfills and the things only testing makes true

A pipeline is finished when running it twice produces the same result as running it once. That is the whole of idempotency, and it is why our own reconciliation job moves each processed file into a migrated/ directory rather than deleting it — you cannot rerun a job against evidence you threw away.

Backfills deserve more suspicion than they get, because a default applied to historic rows is a decision, not a formality. Migrating our stored submissions into Postgres, the historic spam was rescored on the way in and written as held rather than defaulting to clean, which is the difference between a migration that preserves the truth and one that quietly manufactures it.

And a backup is a claim until it has been restored. Ours was verified by restoring into a scratch database and counting rows, not by confirming a file existed on disk with a plausible size.

Where hiring a data engineer is the wrong move

If nobody in the organisation agrees what “active customer” means, a pipeline will encode one person’s answer and make the disagreement permanent. Settle the definitions first; that work is analysis, not engineering.

If the data fits comfortably in the operational database and the actual complaint is that reports are slow, the honest answer is often three indexes and a materialised view, delivered by the application team in a fortnight. A warehouse programme started to solve that is an expensive way to avoid reading a query plan.

And if there is no owner for the alerts, do not build the alerting. Data quality checks that fire into a channel nobody reads are worse than none, because they create the impression of supervision.

What we ask in an interview

How does your pipeline behave on a rerun, and what did you do the first time it did not. What happens when an upstream file arrives late, arrives twice, or arrives with a column removed. How did you find out the last time a number in a report was wrong — from a check, or from a person.

Engineers who have carried a pipeline in production answer with specifics about idempotency, reconciliation and watermarking. The rest describe the happy path, in detail, and at length.

Teams are built for companies in the United States and the Gulf — the UAE, Saudi Arabia, Qatar, Kuwait, Bahrain and Oman. The engineers are in Pune, which matters mostly for the clock. Dubai is ninety minutes behind us and Riyadh two and a half hours, so a Gulf team shares almost the whole working day. New York is nine and a half hours behind, so American engagements run on a written handover and one fixed overlap window rather than on a standing call — a real constraint, and better stated than discovered.

What these engineers do

  • PostgreSQL schema design, indexing, partitioning and query plans under real write load
  • Batch and incremental pipelines in Airflow or Dagster with idempotent reruns
  • dbt models, tests and documented lineage in place of undocumented SQL sprawl
  • Change data capture and warehouse loading into BigQuery, Redshift or Snowflake
  • Data quality checks that block a bad load rather than reporting it the next morning

Delivered AI-first

AI assistance is used for the parts of data work that are structurally repetitive - generating dbt model and test stubs from a source schema, drafting pipeline DAGs, translating stored procedures into versioned SQL, and writing the migration scripts that accompany a schema change. Correctness is established by tests and by reconciliation against source counts, never by the assistant asserting it. Review of schema changes and migrations is unchanged. The measurable effect is throughput per engineer, not fewer reviews.

Built with Data Engineering at Sazinga

These are our own production applications, not client references — which is why the engineers have operated them, not just written them.

Tell us what the Data Engineering work is.

Roughly what it involves, the seniority you need, and when it has to start. We will say what it takes to staff it, or say honestly that we are not the right people for it.

A person reads every enquiry and replies within one working day.