lab

Why Your Backend Is Slow Even When CPU Looks Fine

A failure-oriented walkthrough of connection-pool saturation, database contention, queueing, and the misleading comfort of normal host-level CPU.

A backend can spend most of its time waiting while every CPU dashboard looks reassuring. That makes low CPU a weak argument that capacity is healthy.

This lab establishes a small Spring-style request path backed by PostgreSQL, then introduces three bottlenecks that create similar user-facing latency while requiring very different fixes.

The failure scenario

Assume a service whose median latency is stable but whose p95 moves from 180 ms to 900 ms under load. Application CPU remains below 55%.

The first three hypotheses are:

  1. database connection-pool saturation;
  2. lock or query contention inside the database;
  3. downstream queueing caused by retries.

The important point is that each can produce waiting without producing proportionally high application CPU.

What to instrument

Start with the request path rather than the host:

Layer Signal
HTTP p50 / p95 / p99 latency, request concurrency
Application active threads, blocked time, queue depth
Connection pool active, idle, pending, acquisition time
Database query latency, locks, active sessions, I/O
Retry path retry count, backoff, duplicate concurrency

A useful diagnosis should tell us where time accumulates, not merely which machine looks busy.

A deliberately bad fix

Adding application instances can make the system worse when the scarce resource is database connections. Every new instance contributes another connection pool, increasing pressure on the same database while spreading CPU across more hosts.

This is a recurring systems pattern: horizontal scaling of the wrong layer can amplify contention.

Decision rule

Before adding capacity, identify the resource whose wait time rises with user-visible latency. If no measured wait explains the latency, the diagnosis is incomplete.

Next experiment

The follow-up lab compares three remediations: a larger database, read isolation, and reducing synchronized demand with caching and asynchronous work.

Applied useUseful when an API is persistently slow but the obvious host-level metrics do not identify a bottleneck.