PyTorch Foundation: Six Open-Source Projects Accelerate AI Development, FlexAttention Now 12x Faster
The PyTorch Foundation, a multi-project foundation since April 2025, brings together six open-source projects. PyTorch 2.13 delivers FlexAttention 12x faster on Apple Silicon and up to 4x lower GPU memory usage, with 4,415 commits in Q2. vLLM, DeepSpeed, Ray Data, and Helion also report progress.
This article was generated using artificial intelligence from primary sources.
PyTorch Foundation brings together six open-source projects
The PyTorch Foundation, a multi-project foundation established in April 2025, today brings together six open-source projects: PyTorch, vLLM, DeepSpeed, Ray, Helion, and other components of the machine learning ecosystem. Shared governance enables coordinated development of tools spanning training, serving, and model optimization.
How much faster is FlexAttention in PyTorch 2.13?
The new PyTorch 2.13 release brings FlexAttention, an attention mechanism that is now about 12x faster on Apple Silicon than in previous versions. The new nn.LinearCrossEntropyLoss function reduces required GPU memory by up to 4x, and the second quarter recorded 4,415 commits, significantly more than comparable earlier periods in the project.
Progress at vLLM, DeepSpeed, Ray Data, and Helion
vLLM has established a stable two-week release cycle and announced its first dedicated conference, running August 24-26 at the Ray Summit. DeepSpeed shipped six releases in the first quarter, and its SuperOffload received an honorable mention for best paper at the ASPLOS 2026 conference. Ray and vLLM on AMD MI325X GPUs achieve a 67 percent savings in prefill-decode disaggregation, while the upcoming Ray Data 2.57 announces a completely new high-performance engine. Helion delivers 10x faster autotuning and state-of-the-art results on NVIDIA Blackwell and Google TPU architectures.
Frequently Asked Questions
- What is the PyTorch Foundation?
- The PyTorch Foundation is a multi-project foundation established in April 2025 that brings together six open-source projects from the machine learning ecosystem, including PyTorch, vLLM, DeepSpeed, Ray, and Helion.
- What is FlexAttention, and how much faster is it?
- FlexAttention is an attention mechanism in PyTorch that in version 2.13 runs about 12x faster on Apple Silicon, while the new nn.LinearCrossEntropyLoss function reduces required GPU memory by up to 4x.
Sources
📬 AI news in your inbox
A daily digest built your way — pick topics, sources and cadence. One-click unsubscribe.
Related news
CNCF: Confidential Containers Advances to Incubating Status with 150+ Contributors and a Focus on Confidential AI
NVIDIA: Open-Source Medical Physics Framework Cuts Robot Policy Training from 5 Hours to 2 Minutes
Meta: SAM 3 and DINOv3 open-sourced for the Genesis Mission, segmentation analysis cut from a month to 15 minutes