To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

Imported from official source

AI Classified by Officially

Dylan Lim, Xinyi Li, Sonny Li, Peter Wu, Dan Fu, Simran Arora

40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...

The kernels team at Together recently received access to the NVIDIA Vera Rubin NVL72 platform. We spent the past few days digging through the new ISA and poking the chip with micros. There are many fun new features! We finished adding some functionality into ThunderKittens to write NVFP4 and FP8 GEMMs on Vera Rubin, as well as help fellow kittens explore the stars.

Before diving into what Vera Rubin delivers, we start our journey with a quick refresher of a Blackwell GPU GEMM.

NVIDIA Blackwell architecture’s fifth-generation tensor cores fundamentally changed the GEMM programming model. While NVIDIA Hopper architecture’s wgmma instruction was issued collectively by a warpgroup, Blackwell’s tcgen05 instruction is issued by a single thread, enabling one small producer warp to drive the tensor cores. The accumulator also moved out of registers into Tensor Memory and operands are read directly from shared memory, enabling a single MMA to span two CTAs across two SMs.

To reach competitive performance on Blackwell, our GEMM:

  • Launches threadblock clusters so each CTA pair can share operands by TMA multicast, cutting memory traffic from HBM by half.
  • Specializes warps within clusters: loaders bring A and B into shared memory over TMA, a single MMA warp drives the tensor cores, and a consumer warpgroup carries finished accumulators from tensor memory to HBM.
  • Runs persistently, with one tile's inputs streaming in while the previous tile's outputs are still draining.
  • Through these efforts, we realized the following results.

    Read more about these kernels and their optimizations in our earlier Together blog post or the ThunderKittens 2.0 release!

    This is an extract. The publication continues at the source.

    Read the original at the source: https://www.together.ai/blog/to-infinity-and-beyond-thunderkittens-now-on-nvidia-vera-rubin-nvl72

    Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

    Provenance

    Organization
    Together AI — imported from official source
    Official source
    https://www.together.ai/blog/rss.xml RSS
    Imported
    September 20, 2026 19:52
    Versions
    1 recorded
    Identity
    https://www.together.ai/blog/to-infinity-and-beyond-thunderkittens-now-on-nvidia-vera-r...

    Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.