PyTorch 2.3: User-Defined Triton Kernels in torch.compile, Tensor Parallelism in Distributed
Imported from official source
# PyTorch 2.3 Release notes * Highlights * Backwards Incompatible Changes * Deprecations * New Features * Improvements * Bug fixes * Performance * Documentation # Highlights We are excited to announce the release of PyTorch® 2.3! PyTorch 2.3 offers support for user-defined Triton kernels in torch.compile, allowing for users to migrate their own Triton kernels from eager without experiencing performance complications or graph breaks. As well, Tensor Parallelism improves the experience for training Large Language Models using native PyTorch functions, which has been validated on training runs for 100B parameter models. This release is composed of 3393 commits and 426 contributors since PyTorch 2.2. We want to sincerely thank our dedicated community for your contributions. As always, we encourage you to try these out and report any issues as we improve 2.3. More information about how to get started with the PyTorch 2-series can be found at our [Getting Started](https://pytorch.org/get-started/pytorch-2.0/) page. Stable Beta Prototype Performance Improvements User-defined Triton kernels in torch.compile torch.export adds new API to specify dynamic_shapes Weight-Only-Quantization int...
This version
- Version
- 2 of 2
- Recorded
- September 17, 2026 21:30
- Change
- Imported change
- Content hash
df1754e477deecc9cad92d3a9c156b31- All versions
- Revision history