PyTorch 官方最新动态:Session-Aware Agentic Inference with NVIDIA Dynamo

ADK PyTorch官方 / ADK编译 2026-10-08T18:13:46+00:00 3分钟 166 次浏览
速览导读 / Summary

Agentic workloads change the traffic an inference server sees. Unlike single-turn chat, an agent session can involve a large initial prefill, repeated model calls, and subagents running in parallel, w

官方发布 2026-10-08T18:13:46+00:00 官网即时同步

Key Insights / 核心看点

  • 1 Agentic workloads change the traffic an inference server sees. Unlike single-tur

PyTorch 官方最新动态:Session-Aware Agentic Inference with NVIDIA Dynamo

来源:PyTorch 官方动态 | 发布日期:2026-10-08T18:13:46+00:00

核心更新概览

Agentic workloads change the traffic an inference server sees. Unlike single-turn chat, an agent session can involve a large initial prefill, repeated model calls, and subagents running in parallel, w

详细内容记录

Agentic workloads change the traffic an inference server sees. Unlike single-turn chat, an agent session can involve a large initial prefill, repeated model calls, and subagents running in parallel, with KV cache remaining resident while tools execute between turns. This technical blog post details how NVIDIA Dynamo uses a unified, session level identifier to transform request-level serving infrastructure into a program-aware system, unlocking session-aware routing, shared KV cache indexing, and programmatic cache movement across vLLM and SGLang. Agentic workloads have changed the shape of the traffic an inference server sees. A coding session opens with a large prefill, often tens of thousands of tokens of system prompt, tool definitions, and user-specific guidance, and every turn after that appends to a context the session resends in full. One task fans out into dozens of such calls, alongside short and long lived subagents running in parallel. Most of a session’s wall-clock time passes between tool calls, with its context sitting resident in the KV cache while nothing generates. Serving these well comes down to how much of that context a system keeps resident and reuses, and how many sessions it holds at once while still meeting latency targets This contrast between agentic workloads and that of a standard chatbot is shown in Figure one. Figure 1: Compares the alternating user and model turns of a standard chatbot with an agentic workflow that includes tool calls and tool responses. In the agentic workflow, a single user request can trigger multiple inference calls, with subsequent steps depending on intermediate results. Most open-source serving stacks are still built for the model on top of the figure. They route, admit, and cache each request on its own, with no notion of the harness or the session that ties it to previous turns. Serving this shape optimally is something we have been building towards. In our previous post, Full-Stack Optimizations for Agentic Inference , we shared work across three layers of the NVIDIA Dynamo stack: the router, the inference engine, and the KV cache manager. Over the last few months we’ve unified this work around a single primitive: a common session-level identifier. Session-level ID (often called program or trajectory in papers and academia) turns request-aware infrastructure into program-aware infrastructure, and serves as the foundation for additional session-aware optimizations.

更多技术细节可访问官方原文:https://pytorch.org/blog/session-aware-agentic-inference-with-nvidia-dynamo/。

code · 免费+付费
★ 5.0 · 120评测
P

PyTorch

开源的机器学习库

PyTorch 是开源的机器学习库,主要用在深度学习研究和应用开发,以灵活性、易用性和强大的 GPU 加速功能而闻名。PyTorch 提供动态计算图,支持开发者在运行时动态修改模型结构,非常适合快速开发和实验。PyTorch 支持张量计算、自动微分(torch.autograd)和模块化的神经网络构建(torch.nn)。PyTorch 拥有丰富的社区支持和大量的预训练模型及教程,是学术界和工业界的首选深度学习框架之一。

查看 PyTorch 使用教程与功能