ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

Together AI Version 1 original current

Imported from official source

ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.

This version

Version
1 of 1
Recorded
September 20, 2026 19:52
Change
Initial
Content hash
529631d52632cda216ef8592893d77fb
All versions
Revision history

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.