ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale
AI Classified by Officially
A program-level scheduler that eliminates KV cache thrashing, delivering more than 2x single-node throughput and near-linear multi-node scaling for large-scale agentic inference.
Hao Kang, Ziyang Li, Weili Xu, Xinyu Yang, Yinfang Chen, Junxiong Wang, Beidi Chen, Tushar Krishna, Chenfeng Xu, Simran Arora.
40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...
We introduce ThunderAgent, a system for high throughput agentic inference. By introducing a novel program abstraction for agentic LLM request scheduling, ThunderAgent achieves up to 2.5× higher single-node throughput in our synthetic data generation pipeline, and delivers 2.4× speedup on an 8-node cluster with near-linear throughput scaling with respect to GPU nodes.
→ More than 2× single-node throughput, with roughly 10× lower P50 latency at high concurrency
→ 2.4× speedup on 8 nodes, near-linear scaling from 16 to 64 GPUs
→ Drop-in: one program_id field, OpenAI-compatible, works with your existing engine-level optimizations (like speculative decoding)
ThunderAgent was accepted to ICML 2026 as a Spotlight paper.
This paper was a collaboration between researchers at Georgia Institute of Technology, the University of Illinois Urbana-Champaign, Carnegie Mellon University, and Together AI.
LLMs are increasingly deployed as agents. Systems like Claude Code, Codex, and OpenClaw reason, call tools, read results, and reason again, often for dozens of turns before completing a task. Training these agents requires large scale synthetic data generation, since natural web corpora do not contain multi-turn, tool-using agent trajectories. To generate agentic datasets like our recently released CoderForge, we need to run agentic inference at high concurrency.
This is an extract. The publication continues at the source.
Read the original at the source: https://www.together.ai/blog/thunderagent
Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.
Provenance
- Organization
- Together AI — imported from official source
- Official source
- https://www.together.ai/blog/rss.xml RSS
- Imported
- September 20, 2026 19:52
- Versions
- 1 recorded
- Identity
-
https://www.together.ai/blog/thunderagent