ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)
Imported from official source
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.
This version
- Version
- 1 of 1
- Recorded
- September 20, 2026 19:52
- Change
- Initial
- Content hash
529631d52632cda216ef8592893d77fb- All versions
- Revision history