One Request, End to End · Episode 11

What changes when 10,000 users place orders at once?

Large amounts of client traffic converging on backend application servers
Episode 11Survive the spike
Episode 11 of 12Series roadmap

At 100 orders per minute, the system looks healthy. When 10,000 people try at the same time, the same code starts timing out. The request still follows the same path, but every queue that was almost empty before is now filling up.

Every component has a limit. During a spike, the system has to control how much new work it accepts, decide how long that work may wait, and preserve enough capacity to finish the requests already in progress.

Arrival rate meets service rate

If requests arrive faster than a component completes them, a queue grows.

arrival:  2,000 requests/second
service:  1,500 requests/second
backlog:    500 more requests every second

The queue may be explicit in a broker or hidden in a socket backlog, proxy, event loop, connection pool, lock wait list, or client retry loop. Hidden queues are still queues.

A short queue can smooth out a brief burst. An unlimited queue turns overload into growing memory usage and very long response times. Every waiting line needs a maximum size, a deadline, and a policy for rejecting work the system cannot finish on time.

The first bottleneck moves

Adding application instances removes one limit and increases pressure elsewhere. Ten instances with a pool of 30 database connections can open 300 sessions. Scaling to 50 instances makes that 1,500.

The database may then spend more time managing sessions, scheduling queries, and waiting on locks while completing less useful work. A larger pool can make that problem worse. The better fix may be smaller per-instance pools, a connection pooler, fewer queries per request, controlled concurrency, caching, or additional database capacity.

Capacity planning therefore has to follow the pressure from the application tier into every dependency behind it.

Hot keys serialize traffic

The system may have plenty of total database capacity and one contested row. A flash sale sends thousands of updates to the same inventory record. Row-level correctness turns that record into a serialization point.

Possible responses include:

  • accept that one SKU has finite serialized throughput;
  • reserve inventory in bounded batches closer to workers;
  • partition demand by a safe business key;
  • use an admission queue for the hot item;
  • respond quickly when capacity is exhausted;
  • redesign the product interaction, such as a waiting room or lottery.

Sharding the orders table does not remove the rule around one final unit of a popular item. If only one request can win, the system cannot parallelize that decision without changing how the product allocates inventory.

Retries create more traffic

At high load, latency rises. Clients time out and retry. Each retry adds work to the overloaded system, increasing latency and causing more retries.

If the client, edge, application, and worker all retry three times, one failed operation can turn into many downstream attempts. For each boundary, decide which component owns the retry and give it a fixed budget.

Backoff and jitter spread attempts over time. Idempotency protects write semantics. Neither makes an unavailable dependency available, so overload controls must still reject or defer work.

Backpressure protects the dependency

Backpressure means the part producing work adjusts to what the next component can actually handle.

In an HTTP API, that can mean:

  • limit concurrent calls to payment;
  • stop reading a large upload when downstream processing is full;
  • cap waiting requests;
  • return 429 Too Many Requests or 503 Service Unavailable with a meaningful retry policy;
  • move acceptable work into a durable bounded queue;
  • prioritize checkout completion over noncritical analytics.

Rejecting a request immediately can feel harsh, but during overload it is often better than accepting everything and letting every user wait for a timeout.

Autoscaling arrives after the spike

Autoscaling has a delay. The platform has to observe the signal, decide to scale, start new instances, and wait for them to become ready. If the database or payment quota is already the bottleneck, adding application instances will not help anyway.

Choose scaling signals connected to demand and saturation. CPU may work for CPU-bound services. Queue depth, concurrency, or request rate can be better for I/O-heavy services. Always include maximums so a bad dependency does not cause uncontrolled fleet growth.

Keep enough warm capacity for the ramp time the product cannot tolerate.

Cache only what may be stale

Caching product descriptions can remove repeated reads. Caching the last unit of inventory is much harder because that value decides who can buy it. The cache is another copy of the value, with its own update timing and expiration rules.

Use caches where the product can define acceptable staleness and a recovery path. Protect the database from cache misses with request coalescing and bounded fill concurrency, or one expired popular key can create a stampede.

Test the way traffic actually arrives

A useful load test models:

  • realistic route mix and payload sizes;
  • connection reuse and TLS behavior;
  • gradual ramps, sudden spikes, and sustained plateaus;
  • hot products and skewed customer behavior;
  • downstream latency and failure;
  • retries and client deadlines;
  • enough data for realistic query plans;
  • steady-state recovery after load falls.

An open-loop test keeps sending at the planned arrival rate even when responses slow down, so it exposes overload. A closed-loop test waits for a response before sending the next request, which means the test may send less traffic as latency rises. Use the model that matches the question you are trying to answer.

Test in an environment where safety limits and dependencies are understood. A load test against shared production services can become the outage it was meant to prevent.

Define a capacity envelope

“Handles 10,000 users” is too vague to test. A useful capacity target specifies:

  • requests per second and their mix;
  • concurrent connections;
  • target latency percentiles;
  • acceptable error rate;
  • payload size;
  • data volume and hot-key distribution;
  • dependency latency;
  • durability and consistency requirements;
  • duration of the workload.

After defining those numbers, identify which resource saturates first and what users experience after it reaches that limit.

More traffic does not automatically mean the monolith has to become microservices. It means the capacity assumptions need real numbers. The fix may be one index, a smaller connection pool, a bounded queue, or a sensible limit. I would only add a new service boundary when the measurements show why it is needed.

The final episode puts the full request path back together and shows how pressure or failure in one part affects everything around it.

Sources and further reading