ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)
AI Classified by Officially
The best frontier model solves under a third of 87 real-world problems — but a few generated kernels beat anything publicly available.
Willy Chan, Nathan Paek, Simon Guo, Simran Arora, Daniel Y. Fu
40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...
LLMs have gotten surprisingly good at writing GPU kernels[1][2][3], but almost all current benchmarks measuring that progress are single-GPU. In production, communication is often the bottleneck: communication overhead can account for over 20% of inference latency[4], and that gap keeps widening as compute scales faster than interconnect bandwidth.
ParallelKernelBench (PKB) offers a benchmark and evaluation framework for multi-GPU kernel generation and includes 87 problems from real codebases where the task is replacing PyTorch + NCCL with a CUDA kernel that moves data directly over NVLink. We tested frontier coding models such as GPT-5.5, Gemini 3 Pro, Opus 4.7, and others. The evaluation revealed significant performance gaps across the board: under a third of problems were solved correctly, and fewer than a quarter of those beat the naive baseline.
We'll cover why they fail, what the patterns look like, and a few cases where models surprisingly produced kernels faster than anything publicly available, including one for NVIDIA NeMo-RL's GRPO training loop, which has no prior optimized public reference.
Why multi-GPU is different from single-GPU kernel generation
LLMs have made progress on GPU kernel generation, but that progress has mostly been measured on a single GPU. Production AI workloads no longer fit that frame: they span multiple GPUs, and performance is increasingly shaped by communication rather than just local compute and memory. That shift makes multi-GPU kernel generation a different problem in three ways:
This is an extract. The publication continues at the source.
Read the original at the source: https://www.together.ai/blog/parallelkernelbench
Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.
Provenance
- Organization
- Together AI — imported from official source
- Official source
- https://www.together.ai/blog/rss.xml RSS
- Imported
- September 20, 2026 19:52
- Versions
- 1 recorded
- Identity
https://www.together.ai/blog/parallelkernelbench