Skip to content

Free lesson · about 10 minutes

Ready, warm, useful

A green readiness check says a replica may take traffic. It doesn't say the first user gets a fast answer. This lesson separates the three things that get blurred together after a restart: ready, warm, and useful.

No sign-up, no tracking, nothing stored. Every number on this page is an invented teaching constant, not a hardware measurement.

The misconception

“Ready means the first request is already fast.” In Kubernetes, a readiness probe decides whether a container receives traffic. It passes when whatever the probe checks passes. For a model server that's usually “the weights are loaded and the port answers”, which is a long way from “a user's request finishes quickly”.

Three milestones are worth keeping apart:

  • Ready: the probe passes, so traffic is allowed in. A server state.
  • Warm: the work a request can reuse already exists. For LLM serving, the clearest example is a cached prompt prefix, which lets the server skip part of prefill.
  • Useful: a person got an answer they could act on. A user outcome, and the one people actually feel.

Try a deliberately simple model

One replica, already ready at time 0. It serves two requests at once. Each request pays a prefill stage and a decode stage. A cold prefix costs 8 toy units of prefill; a warm one costs 2, because only the part that isn't cached is computed. Decode costs 4 either way. Pick a scenario.

Scenario

Request timeline
Show the numbers as a table
RequestQueuePrefillDecodeOutcome atOutcome
Constants and every scenario's timeline, for your own tests.

What the three scenarios show

One request. The replica is ready in both states. The first request takes 12 units cold and 6 warm. Readiness was identical, so readiness can't be what made the difference.

A burst after restart. Requests 3 and 4 take 18 units, but only 6 of those are work. They spent 12 waiting for a slot, and by the time they ran, the prefix was warm. If you only record total latency, you can't tell whether to fix the queue or the prefill. And the first wave paid the cold prefill twice, because nothing in this toy coalesces identical work.

Degraded mode. Admitting two and telling the other two “busy, retry” at once doesn't finish more work by time 12. What it changes is the longest stretch anyone sat in a queue without being told anything: 12 units down to 1. It still has to be honest. A degraded response that pretends to be a full answer is worse than a slow one.

Check your reasoning

Three questions. Each answer explains itself, right or wrong.

1. A replica is ready. Its prefix cache is cold. What can you conclude?
2. After a change, pods become ready sooner, but first-request latency is unchanged. What did you observe?
3. In the burst, request 3 took 18 units and request 1 took 12. Which observation separates queue delay from prefill delay?

What to measure on a real system

  1. Record the readiness time and the first request after readiness as two separate numbers.
  2. Label each early request with what was already present: image, weights, and prefix state.
  3. Split each request's time into queue wait, prefill and decode. Totals alone can't tell you which stage to fix.
  4. Credit an optimization only with the stage it changed. A faster pod start isn't a prefill fix.
  5. Rehearse a cold restart under a burst, not just one polite request.
  6. Write down what callers are told while the system is degraded, and check that it's true.

What this lesson does not show

  • No GPU speedup. The 8 and 2 are invented so the arithmetic is easy to follow. Real prefix-cache gains depend on the serving engine, how much of the prompt matches, and the workload.
  • No real queueing model. Two fixed slots and four simultaneous arrivals are a teaching shape, not a prediction of any system's latency.
  • Nothing about cost. Who pays for a cached prefix is a separate decision, and it's the subject of Who pays for the KV cache?. Latency and attributed cost are different outputs; keep them on different lines.

Questions or feedback?
I reply to every note.

Say hi