🟢 📦 Open Source Published: · 2 min read ·

PyTorch Foundation: Six Open-Source Projects Accelerate AI Development, FlexAttention Now 12x Faster

PyTorch logo alongside a visualization of a neural network graph and GPU chips

The PyTorch Foundation, a multi-project foundation since April 2025, brings together six open-source projects. PyTorch 2.13 delivers FlexAttention 12x faster on Apple Silicon and up to 4x lower GPU memory usage, with 4,415 commits in Q2. vLLM, DeepSpeed, Ray Data, and Helion also report progress.

🤖

This article was generated using artificial intelligence from primary sources.

PyTorch Foundation brings together six open-source projects

The PyTorch Foundation, a multi-project foundation established in April 2025, today brings together six open-source projects: PyTorch, vLLM, DeepSpeed, Ray, Helion, and other components of the machine learning ecosystem. Shared governance enables coordinated development of tools spanning training, serving, and model optimization.

How much faster is FlexAttention in PyTorch 2.13?

The new PyTorch 2.13 release brings FlexAttention, an attention mechanism that is now about 12x faster on Apple Silicon than in previous versions. The new nn.LinearCrossEntropyLoss function reduces required GPU memory by up to 4x, and the second quarter recorded 4,415 commits, significantly more than comparable earlier periods in the project.

Progress at vLLM, DeepSpeed, Ray Data, and Helion

vLLM has established a stable two-week release cycle and announced its first dedicated conference, running August 24-26 at the Ray Summit. DeepSpeed shipped six releases in the first quarter, and its SuperOffload received an honorable mention for best paper at the ASPLOS 2026 conference. Ray and vLLM on AMD MI325X GPUs achieve a 67 percent savings in prefill-decode disaggregation, while the upcoming Ray Data 2.57 announces a completely new high-performance engine. Helion delivers 10x faster autotuning and state-of-the-art results on NVIDIA Blackwell and Google TPU architectures.

Frequently Asked Questions

What is the PyTorch Foundation?
The PyTorch Foundation is a multi-project foundation established in April 2025 that brings together six open-source projects from the machine learning ecosystem, including PyTorch, vLLM, DeepSpeed, Ray, and Helion.
What is FlexAttention, and how much faster is it?
FlexAttention is an attention mechanism in PyTorch that in version 2.13 runs about 12x faster on Apple Silicon, while the new nn.LinearCrossEntropyLoss function reduces required GPU memory by up to 4x.

📬 AI news in your inbox

A daily digest built your way — pick topics, sources and cadence. One-click unsubscribe.