Why Your Backend Is Slow Even When CPU Looks Fine
A failure-oriented walkthrough of connection-pool saturation, database contention, queueing, and the misleading comfort of normal host-level CPU.
A backend can spend most of its time waiting while every CPU dashboard looks reassuring. That makes low CPU a weak argument that capacity is healthy.
This lab establishes a small Spring-style request path backed by PostgreSQL, then introduces three bottlenecks that create similar user-facing latency while requiring very different fixes.
The failure scenario
Assume a service whose median latency is stable but whose p95 moves from 180 ms to 900 ms under load. Application CPU remains below 55%.
The first three hypotheses are:
- database connection-pool saturation;
- lock or query contention inside the database;
- downstream queueing caused by retries.
The important point is that each can produce waiting without producing proportionally high application CPU.
What to instrument
Start with the request path rather than the host:
| Layer | Signal |
|---|---|
| HTTP | p50 / p95 / p99 latency, request concurrency |
| Application | active threads, blocked time, queue depth |
| Connection pool | active, idle, pending, acquisition time |
| Database | query latency, locks, active sessions, I/O |
| Retry path | retry count, backoff, duplicate concurrency |
A useful diagnosis should tell us where time accumulates, not merely which machine looks busy.
A deliberately bad fix
Adding application instances can make the system worse when the scarce resource is database connections. Every new instance contributes another connection pool, increasing pressure on the same database while spreading CPU across more hosts.
This is a recurring systems pattern: horizontal scaling of the wrong layer can amplify contention.
Decision rule
Before adding capacity, identify the resource whose wait time rises with user-visible latency. If no measured wait explains the latency, the diagnosis is incomplete.
Next experiment
The follow-up lab compares three remediations: a larger database, read isolation, and reducing synchronized demand with caching and asynchronous work.