kv-cache 3 Where gpu_memory_utilization Actually Goes: vLLM's Memory Budget, Line by Line Jun 4, 2026 SGLang Deep Dive: How Engine.generate() Boots and Runs Apr 28, 2026 vLLM v1 Deep Dive: How LLM(...) and the Server Boot and Generate Apr 27, 2026