What does it actually take to run agentic workloads at scale? ⚡Agents push token consumption, conte...
NVIDIA AI(@NVIDIAAI)150 字 (约 1 分钟)
52
NVIDIA 宣称其 Vera Rubin 平台通过软硬协同设计,支持高吞吐、长上下文、低延迟的智能体(agent)推理负载,实测达 400+ tokens/sec/user。
入选理由:智能体工作负载对 token 消耗、上下文长度和延迟提出极端要求
FeaturedTweet#NVIDIA#AI Infra#LLM Agents#Hardware Acceleration中英混合
