All posts

What I actually do as a Kafka Connect engineer

· 4 min read

When people ask what I do, “I build Kafka connectors” usually gets a polite nod. So here’s the longer version, for friends, family and anyone thinking about this kind of work.

First, what is Kafka?

Think of Kafka as a very fast, very reliable conveyor belt for data. One part of a company puts events on the belt: an order was placed, a payment went through, a sensor took a reading. Other parts pick those events up and do something with them, often within milliseconds.

Kafka is great at moving data around inside a company. The hard part is getting data onto the belt from everywhere it already lives, and off the belt into everywhere it needs to go.

Where connectors come in

That’s what a connector does. There are two kinds:

  • A source connector reads from an outside system and puts the data into Kafka.
  • A sink connector takes data from Kafka and writes it to an outside system.

Without them, every team would write and babysit its own little sync script. With them, you write a few lines of configuration and Kafka Connect does the rest: it runs the connector, spreads the work across machines, keeps track of what’s been copied, and picks up where it left off if something crashes.

What I work on

My team looks after connectors for object storage, which means Amazon S3, Azure Blob Storage and Google Cloud Storage, plus the HTTP connectors that can talk to almost any web API.

Object storage is where a lot of data ends up for the long term: cheap, durable, and easy to query later. A typical setup looks like this:

{
  "connector.class": "io.confluent.connect.s3.S3SinkConnector",
  "topics": "orders",
  "s3.bucket.name": "company-data-lake",
  "s3.region": "ap-south-1",
  "format.class": "io.confluent.connect.s3.format.json.JsonFormat",
  "flush.size": "1000"
}

That’s it. Every 1,000 orders, a new file lands in the bucket. The connector handles retries, batching, file naming and keeping track of exactly which records it has already written.

Why it’s harder than it looks

Copying data sounds simple until you remember that it has to be correct, all the time, at scale. A few things I think about every day:

  • Nothing lost, nothing doubled. If a machine restarts halfway through writing a file, the data must not go missing and must not be written twice.
  • Other people’s systems fail. Cloud APIs time out, rate-limit you or return errors that mean three different things. A good connector knows when to retry, when to wait and when to stop and tell someone.
  • Data is messy. Formats change, fields go missing and someone sends a 50 MB record. The connector has to cope without falling over.
  • It runs for years. A connector isn’t a script you run once. Once it’s set up, it runs nonstop, so small inefficiencies become real costs.
A typical day, roughly

Some code, some reviews, some reading logs. A good chunk of the job is reproducing a problem on my own machine: spinning up Kafka and a connector locally, feeding it the kind of data that broke it, and watching what happens. Once I can make a bug happen on demand, fixing it is usually the easy part.

What I like about it

It’s real plumbing. When it works, nobody notices, and that’s the point. When a company moves billions of records a day without thinking about it, some of that quiet reliability is code my team wrote.

It also keeps me broad. In one week I might read up on how S3 handles uploads, debug a Java thread that won’t stop, and work out why an HTTP API returns a different error on Tuesdays.

If you want to try it

The Kafka Connect documentation is the best place to start. Run Kafka locally, set up a simple file source connector, and watch your data appear in a topic. It’s a satisfying first step.

I’ll write more about the specifics in future posts, like how connectors avoid writing duplicates and what makes a good retry strategy.

Was this useful?

Mohammed Ashfaq Ahmed

Senior Software Engineer at Confluent, based in Hyderabad. I write about building software and what I learn along the way. More about me