Tensors and neural networks in Python with strong hardware acceleration. PyTorch is an open source project at the Linux Foundation. #PyTorchFoundation

ERRC: Entropy-Reinvested Residual Correction for Tensor-Parallel LLM Communication PyTorch Foundation Ambassador Abdulsalam Bande will present a poster on ERRC (Entropy-Reinvested Residual Correction for Tensor-Parallel LLM Communication) at PyTorch Conference North America 2026. ERRC focuses on optimizing GPU data exchange during LLM inference. By compressing inter-GPU communication and using the saved bandwidth to mitigate quantization errors, ERRC makes multi-GPU inference significantly more efficient. Register and join us to San Jose, October 20-21: hubs.la/Q04v88dJ0 #PyTorchCon @abdulsalambande
7
3
37
15,707
Modern AI workloads rely on a growing range of specialized GPU kernels, and no single kernel generation backend is optimal across every operator. At #PyTorchCon North America 2026, Liz Li and Jiahui Cao (@AMD) will present “Extending TorchInductor with FlyDSL: A New MLIR-Native Backend for High-Performance GEMMs.” Liz and Jiahui will share how FlyDSL, AMD’s Python-native and MLIR-based GPU kernel DSL, integrates with TorchInductor’s compilation and autotuning pipeline while preserving the torch.compile user experience. The session will also present performance results for transformer training and inference on AMD Instinct GPUs, comparing Triton and FlyDSL implementations. Register for PyTorch Conference North America 2026: hubs.la/Q04v4SL60
4
5
45
16,575
Helion, PyTorch's kernel DSL, is showing what autotuned high-level kernels can do for production inference. In this post, we integrate Helion into @vllm_project linear backend and show how a single Helion GEMM implementation can cover multiple algorithmic variants — Standard GEMM, Split-K, and Swap-AB — with the best variant and config selected automatically per shape. On @nvidia Hopper GPUs, combining per-shape tuning with hybrid dispatch outperforms vLLM's default CUTLASS and DeepGEMM backends across the evaluated models, with consistent end-to-end gains and more than 10% throughput improvement on some workloads. Read our latest blog: pytorch.org/blog/building-a-… ✍️ Sean Chen (@RedHat) and Shangdi Yu (PyTorch, @Meta Platforms)
10
13
111
16,864
NEW! Take the PyTorch Certified Associate (PTCA) Certification Pathway which combines focused learning modules & the PyTorch Certified Associate (PTCA) certification exam together in one structured experience. The pathway includes 4 self-paced learning modules plus the PTCA certification exam with a focus on: PyTorch Fundamentals Data Handling in PyTorch PyTorch Model Development Optimizing PyTorch Together, the modules cover core skills across the same key areas measured by the PTCA certification, from working with tensors and devices to preparing data, building neural networks, and improving model performance. Learn more: pytorch.org/blog/new-pathway…
1
34
15,533
New keynote alert! 🎤 Rachad Alao, VP of Engineering at @Cohere, joins #PyTorchCon North America to explore how open science can help AI foundation models advance with transparency, security, speed, and scale. 📅 October 20-21 📍 San Jose Save your seat: bit.ly/4sh3DSw
2
5
21
11,181
Our latest blog presents work on Jagged Flash Attention (JFA) — the attention kernel behind @Meta's Generative Ads Model (GEM) — on @NVIDIA Blackwell (B200), built with TLX (Triton Low-level Extensions), which add explicit, hardware-aware control on top of Triton's high-level, tile-based programming model. The TLX attention kernel is about 3.2K lines of concise Triton-level code — roughly 3× less than the ~10K-line CuteDSL kernels of the state-of-the-art FlashAttention-4 (FA4), and it outperforms FA4 on the jagged shapes that matter for GEM — by ~13% on the forward pass and ~50% on the backward pass. Read it here: pytorch.org/blog/optimizing-… ✍️ Han Xu, Jacky Zhou, Jackie (Jiaqi) Xu, Hongtao Yu, Peng Chen (Dev Infra), Darren Liu, Dev (Devashish) Shankar (@dev17392), Max Leung, Nick Riasanovsky, Hao Yan, Manman Ren, Yuanwei (Kevin) Fang
8
29
260
18,737
PyTorch retweeted
We’ll be in San Jose for @PyTorch Conference North America Oct 20 - 21. 🎤 Talks on Triton, MoE fine-tuning + model customization 🤖 Booth demos + challenges 🍻 AI Infra Night with @dstackai, @SGLang (@lmsysorg) Register: 🔗 luma.com/6j69xyal #PyTorchCon ​​#PyTorchFoundation
4
4
10
3,940
Want to hear more about new vLLM and SGLang engines built on a new PyTorch-native TPU backend, TorchTPU? Find Qi Zhou of @Google, Colin Taylor and Angela Yi of @Meta at PyTorch Conference North America in San Jose where they will discuss how they focused on making the existing @sgl_project & @vllm_project infrastructure work natively on TPU, preserving the schedulers, batching systems, OpenAI-compatible APIs, and torch.compile workflows already familiar to GPU users. Because the serving engines remain upstream, new models and features can run on TPU without requiring separate TPU-specific implementations. Learn more about how it works across the stacks & what they learned running this in production, including compile time, device placement under tracing, KV cache layout, and cross-chip collectives. Register for #PyTorchCon NA now: hubs.la/Q04v4SL60
9
11
82
16,586
Your next favorite #PyTorchCon North America session might take only 10 minutes. ⏱️ Visit the Demo Theater in the Community Expo for compact demos with plenty of compute. Schedule: bit.ly/4hb0ekq Join us Oct 20-21 in San Jose: bit.ly/4sh3DSw
1
18
14,508
What You Cannot Profile, You Cannot Optimize: Learning to Read PyTorch Traces In his talk at PyTorch Conference North America 2026, Suvaditya Mukherjee from @huggingface will share practical guidance on how to start using the PyTorch Profiler on your models to extract maximum performance efficiency. Whether you are aiming to identify performance bottlenecks or optimize model execution, this talk will help you get the most out of your hardware. Join the open source AI community in San Jose, October 20-21: hubs.la/Q04v4SL60 #PyTorchCon @halcyonrayes
3
3
45
16,162
Enabling PyTorch to train any model on any chip in any cloud for any agent is a ridiculously ambitious goal, and these open source community efforts to build CI for testing accelerators are key. Open source is not only the best way to solve this problem, it’s the only way.
Keeping out-of-tree accelerators aligned with a fast-moving ecosystem is no small task. In a new community blog, contributors from @IBM and @RedHat share how they built Torch Spyre on top of PyTorch's Cross-Repository CI Relay (CRCR) to tackle compatibility testing at scale. In the blog they also share testing and workflow patterns that could benefit accelerator developers across the ecosystem. Read the full technical deep dive here: bit.ly/46XkjGi @mehant1
4
12
5,193
The faster the information can flow, the faster the progress can be made, and that’s exactly where open source really shines.” At PyTorch Conference China 2026, PyTorch Foundation Executive Director @sparkycollier drove home the value of open source as AI hardware, model architectures, and inference engines evolve rapidly. In his keynote, “Building Frontier Intelligence in the Open,” Mark describes the challenge as one of coordination across hardware, models, and inference, with the flow of information as a potential bottleneck. As a concrete example, Mark discusses @Shopify's use of agents in production, where failures can feed back into training with PyTorch and serving with @vllm_project as the model improves. Mark's full keynote from Shanghai: piped.video/DAS-mrer7Os?si=BvoU… Register for PyTorch Conference North America, October 20–21 in San Jose: hubs.la/Q04qPqdB0
4
3
23
11,894
As more teams scale open source AI on secure, governed infrastructure, Ray has become critical across the full AI lifecycle, from data curation to production serving and reinforcement learning. Ahead of PyTorch Conference North America 2026 in San Jose, we put together a guide highlighting key Ray-focused sessions on the schedule. Featured speakers include: Anyscale /@databricks / @UCBerkeley : Ion Stoica @anyscalecompute : Eric Tang, Sumanth Hedge, Josh Lee, Mengjin Yan @Google : Ankita Luthra, Trinadh Kotturu, Jago Macleod @LinkedIn : Tao Huang, Tommy Li @Pinterest : Gaurav Arora, Shunyao Li, Eric Wang @Uber : Ke Chen, Peng Zhang, Xandra Zhu University of Minnesota (@UMNews) : Arun Sharma Explore engineering takeaways from these technical sessions where you will learn about unified AI orchestration with Kubernetes, elastic training stacks, and scalable RL. Read the full guide and map out your schedule: bit.ly/4ysjCku Register for PyTorch Conference North America: bit.ly/464FTbp
4
8
29
13,736
We are especially looking forward to Ion Stoica's keynote on Evolving Ray and Kubernetes Together for the AI Era @istoica05
1
2,287
Planning a Bay Area meetup, hackathon, showcase, or open source AI gathering? Make it part of #OpenSourceAIWeek, October 16-25, anchored by #PyTorchCon North America and #AGNTCon + #MCPCon North America in San Jose. Submit by October 15: bit.ly/4ipYZ39
4
1
24
14,344
Keeping out-of-tree accelerators aligned with a fast-moving ecosystem is no small task. In a new community blog, contributors from @IBM and @RedHat share how they built Torch Spyre on top of PyTorch's Cross-Repository CI Relay (CRCR) to tackle compatibility testing at scale. In the blog they also share testing and workflow patterns that could benefit accelerator developers across the ecosystem. Read the full technical deep dive here: bit.ly/46XkjGi @mehant1
5
5
56
17,004
All PyTorch Conference China 2026 sessions are now available to watch on our YouTube channel. PyTorch Conference China 2026 brought the PyTorch community together in Shanghai on September 8–9 alongside KubeCon + CloudNativeCon and OpenInfra Summit, following sponsor-hosted co-located events on September 7. Technical discussions spanned models, frameworks, distributed training, inference, hardware, cloud native infrastructure, and agents. “Open Source for the AI Era” framed work across those layers. Across co-located sessions, keynotes, technical demonstrations, a PyTorch Foundation press conference, community meetings, and conversations at the PyTorch booth, the program covered hardware adaptation, training and serving, open infrastructure, accelerator integration, and collaboration across open source communities. Watch the full PyTorch Conference China 2026 session playlist: piped.video/playlist?list=PL… Next up, #PyTorchCon North America comes to San Jose October 20–21. Register today: hubs.la/Q04tBgv_0
3
6
33
16,142
PyTorch retweeted
We're excited to have a strong vLLM presence at #PyTorchCon North America this year. 🔷 @simon_mo_ , vLLM core maintainer and Co-founder / CEO of @inferact, will give the keynote on scaling open frontier inference infrastructure 🔷 vLLM core maintainer @nickhill33 and George Novack will be at the "Meet the Developers of vLLM" 🔷 vLLM maintainers @tms_jr, Lucas Wilkinson, and Joseph Groenenboom from @RedHat_AI will be giving talks on Agentic Inference and Attention in vLLM. Thank you to @PyTorch for highlighting the community's work!
vLLM (@vllm_project) is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North America, you’ll find vLLM-related work throughout the program, from Simon Mo’s keynote to technical sessions and posters covering LLM inference and serving. @simon_mo_ will present the keynote “vLLM Update: Scaling Open Frontier Inference Infrastructure,” covering improvements to vLLM’s core architecture and major optimizations in KV cache management and GPU kernels. Across #PyTorchCon, technical sessions dig into vLLM-related inference, including attention, KV cache management and transfer, disaggregated serving, elastic expert parallelism, scaling across hardware without forks, and more. "Meet the Developers of vLLM" will feature George Novack and Nick Hill. Posters cover additional vLLM-related work on custom accelerators, Trainium, expert parallelism and RDMA KV-cache transfer, multi-chip KV-cache transfer, tier-aware routing, and more. Explore the vLLM sessions: events.linuxfoundation.org/p… PyTorch conferences are the open source AI community’s town square, where what’s next gets decided. Register for PyTorch Conference North America, October 20–21 in San Jose: hubs.la/Q04vK-VQ0
4
12
75
13,570
How does the PyTorch Ambassador Program work, and how can you position yourself to become a PyTorch Ambassador? At #PyTorchCon North America 2026, Sahdev Zala (IBM) will present “Inside the PyTorch Ambassador Model: How Ambassadors Grow PyTorch Adoption & How You Can Become One.” Sahdev will cover what ambassadors have built, the challenges they have encountered, the results they have produced, and ways to engage more deeply with the PyTorch community through advocacy, education, and contribution. Register for PyTorch Conference North America 2026: hubs.la/Q04v4SL60
1
7
33
15,958