Let's talk

Integration & Eventing

Hire Apache Kafka Engineers

Kafka engineers who treat topic design, ordering and replay as decisions to be made rather than defaults to be inherited.

Most Kafka problems are decided on the day the topic is created

Most Kafka problems are decided on the day the topic is created and discovered a year later. Partition count and partition key together determine what is ordered relative to what — and ordering is a business question, not an infrastructure one. If events for one order can land on three partitions, they can be processed out of sequence, and a cancellation can be applied before the booking it cancels.

Kafka guarantees order within a partition and makes no promise across partitions. That single sentence is the whole design constraint, and choosing the partition key is therefore choosing which sequences the business is allowed to rely on.

The related trap is partition count, because it can be increased and not decreased, and increasing it rehashes keys onto different partitions. A topic that is repartitioned in place loses the ordering guarantee for keys in flight at the moment of the change.

Delivery semantics and the exactly-once misunderstanding

The second decision is delivery semantics. At-least-once is the sane default and it means consumers must be idempotent, because they will see the same event twice. Teams that assume exactly-once because the documentation mentions it, without enabling or understanding what it covers, end up with duplicates in a downstream system and no idea where they entered.

Exactly-once in Kafka covers reads and writes that stay inside Kafka within a transaction. The moment a consumer writes to a database, calls an API or sends an email, the guarantee stops at the boundary and the consumer is responsible for the rest. Most real consumers cross that boundary on their first line of useful work.

Schemas are what let teams deploy independently

Schemas are the third, and the one that stops teams being able to deploy independently. Without a registry and an agreed compatibility mode, adding a field becomes a coordinated release across every consumer — which is precisely the coupling the event spine was introduced to remove.

Backward compatibility lets consumers upgrade after producers; forward compatibility lets them upgrade before. Choosing one and writing it down is more valuable than the choice itself, because the failure mode is not a bad choice, it is a different unwritten assumption in each team.

Operating a stream, which is where the lessons come from

Then there is operating it. Consumer lag that nobody watches, a rebalance storm caused by a slow consumer inside the poll interval, retention set to a week when the replay you eventually need is from a month ago. We staff people who have been on call for a stream, because that is where these lessons come from.

Lag is the metric that matters and it is frequently monitored as an average across a consumer group, which hides the case that actually causes incidents: one partition falling behind because one key is hot. Per-partition lag is the alert worth having.

Replay as a supported operation rather than an emergency

Retention and compaction are usually inherited from a default and are actually a statement about what the organisation can recover from. If a downstream system can be rebuilt from the topic, the topic is the system of record for that data and its retention should say so.

Replay is worth designing for before it is needed: consumers that can be pointed at an earlier offset without side effects, sinks that merge rather than append, and a way to run a replay without the live consumer group. Teams that have never rehearsed it discover during an incident that replay means re-sending every notification the topic ever produced.

When you do not need Kafka

Two services exchanging a few thousand messages a day do not need a broker with partitions, consumer groups and a schema registry. A queue, or an API with a retry, is less to operate and easier to reason about.

Kafka is also the wrong reach when the requirement is really a shared database. Streaming data into three systems so that all three can answer the same question is often more expensive than letting them query one place, and it introduces three copies that can disagree.

The case for Kafka is genuine fan-out to independent consumers, volumes that a queue would struggle with, or a need to replay history to rebuild a downstream system. Where none of those applies, the operational cost arrives immediately and the benefit does not.

Teams are built for companies in the United States and the Gulf — the UAE, Saudi Arabia, Qatar, Kuwait, Bahrain and Oman. The engineers are in Pune, which matters mostly for the clock. Dubai is ninety minutes behind us and Riyadh two and a half hours, so a Gulf team shares almost the whole working day. New York is nine and a half hours behind, so American engagements run on a written handover and one fixed overlap window rather than on a standing call — a real constraint, and better stated than discovered.

What these engineers do

  • Topic, partition and key design driven by the ordering the business actually needs
  • Consumer groups, offsets, rebalancing and what happens when a consumer dies mid-batch
  • Schema Registry and the compatibility rules that keep producers and consumers deployable
  • Kafka Connect, Streams and change data capture into and out of existing systems
  • Retention, compaction and replay as a supported operation rather than an emergency

Delivered AI-first

Engineers use AI assistance to map existing topics and consumers, draft schema and connector configuration, and generate test harnesses for consumer failure paths. Decisions about ordering, keys and delivery semantics are made by a person who can explain them.

Tell us what the Apache Kafka work is.

Roughly what it involves, the seniority you need, and when it has to start. We will say what it takes to staff it, or say honestly that we are not the right people for it.

A person reads every enquiry and replies within one working day.