PagedAttention Made KV-Cache Feel Like a Memory Allocator

My first GPU-backed run used vLLM’s offline LLM.generate API. The second run put vLLM behind its OpenAI-compatible server and measured streaming TTFT and TPOT. Those experiments showed the runtime from a caller’s point of view. They did not show how serving engines manage memory. To understand the memory problem, I first needed to follow one request through generation: how the model reads its prompt, produces each new token, and retains useful state from earlier tokens in that same request....

September 27, 2026 · Sai Boorlagadda