PagedAttention Made KV-Cache Feel Like a Memory Allocator

My first GPU-backed run used vLLM’s offline LLM.generate API. The second run put vLLM behind its OpenAI-compatible server and measured streaming TTFT and TPOT. Those experiments showed the runtime from a caller’s point of view. They did not show how serving engines manage memory. To understand the memory problem, I first needed to follow one request through generation: how the model reads its prompt, produces each new token, and retains useful state from earlier tokens in that same request....

September 27, 2026 · Sai Boorlagadda

My First GPU-Backed LLM Inference Run

I have been reading plenty about inference systems: KV cache, PagedAttention, prefill, decode, batching, and disaggregated serving. But reading about the machinery is different from running it and watching the machine do work. This post is the first in the series that closes that gap. The goal was intentionally small: run one open-source model on a GPU, record what environment I actually got, generate 100 tokens, and capture enough numbers to make the next experiment less hand-wavy....

August 29, 2026 · Sai Boorlagadda