PyTorch 官方最新动态:Session-Aware Agentic Inference with NVIDIA Dynamo
来源:PyTorch 官方动态 | 发布日期:2026-10-08T18:13:46+00:00
核心更新概览
Agentic workloads change the traffic an inference server sees. Unlike single-turn chat, an agent session can involve a large initial prefill, repeated model calls, and subagents running in parallel, w
详细内容记录
Agentic workloads change the traffic an inference server sees. Unlike single-turn chat, an agent session can involve a large initial prefill, repeated model calls, and subagents running in parallel, with KV cache remaining resident while tools execute between turns. This technical blog post details how NVIDIA Dynamo uses a unified, session level identifier to transform request-level serving infrastructure into a program-aware system, unlocking session-aware routing, shared KV cache indexing, and programmatic cache movement across vLLM and SGLang. Agentic workloads have changed the shape of the traffic an inference server sees. A coding session opens with a large prefill, often tens of thousands of tokens of system prompt, tool definitions, and user-specific guidance, and every turn after that appends to a context the session resends in full. One task fans out into dozens of such calls, alongside short and long lived subagents running in parallel. Most of a session’s wall-clock time passes between tool calls, with its context sitting resident in the KV cache while nothing generates. Serving these well comes down to how much of that context a system keeps resident and reuses, and how many sessions it holds at once while still meeting latency targets This contrast between agentic workloads and that of a standard chatbot is shown in Figure one. Figure 1: Compares the alternating user and model turns of a standard chatbot with an agentic workflow that includes tool calls and tool responses. In the agentic workflow, a single user request can trigger multiple inference calls, with subsequent steps depending on intermediate results. Most open-source serving stacks are still built for the model on top of the figure. They route, admit, and cache each request on its own, with no notion of the harness or the session that ties it to previous turns. Serving this shape optimally is something we have been building towards. In our previous post, Full-Stack Optimizations for Agentic Inference , we shared work across three layers of the NVIDIA Dynamo stack: the router, the inference engine, and the KV cache manager. Over the last few months we’ve unified this work around a single primitive: a common session-level identifier. Session-level ID (often called program or trajectory in papers and academia) turns request-aware infrastructure into program-aware infrastructure, and serves as the foundation for additional session-aware optimizations.
更多技术细节可访问官方原文:https://pytorch.org/blog/session-aware-agentic-inference-with-nvidia-dynamo/。