Inference Infrastructure16 min
Continuous Batching: Serving Many Requests on One GPU
Post I-00 traced a single request through the inference pipeline: prefill processed all input tokens in parallel, decode generated output tokens one at a time, and the KV cache grew with every step. At the end of that trace, we noted that 49 other agents were submitting queries at roughly the sam...
Huang Tzu Lin