paged-attention 2 Where gpu_memory_utilization Actually Goes: vLLM's Memory Budget, Line by Line Jun 4, 2026 vLLM v1 Deep Dive: How LLM(...) and the Server Boot and Generate Apr 27, 2026