lab

Kafka Under Load: When More Partitions Stop Helping

Partition count is a scaling mechanism, not a universal cure. This lab separates broker capacity, key skew, consumer concurrency, and rebalance overhead.

When Kafka consumer lag rises, “add partitions” is an attractive response because it is visible, familiar, and sometimes correct. It is also easy to apply to the wrong bottleneck.

Four different limits

A Kafka pipeline can stop scaling because of broker throughput, a hot partition, insufficient consumer concurrency, or work performed downstream of the consumer.

Those limits require different interventions. Adding partitions only helps when additional parallelism can actually be consumed and the key distribution allows work to spread.

Experiment

Generate two workloads with the same total event rate. The first distributes keys uniformly. The second sends 45% of events to 3% of keys.

Increase partition count in both workloads while recording per-partition throughput, consumer lag, rebalance time, and downstream processing latency.

The uniform workload should gain useful parallelism longer. The skewed workload reaches a point where new partitions exist but the hottest keys remain concentrated.

Operational takeaway

Treat lag as a symptom. First identify whether the waiting work is concentrated by partition, consumer, broker, or downstream dependency. Then scale the constrained dimension.

Applied useUseful when consumer lag rises despite adding partitions or brokers, especially when traffic is uneven across keys.