+1 to same compiler flow on our chips as mainline PyTorch. At Meta, we don't have a separate copy of PyTorch or Dynamo/Inductor. The source of truth is Github and everything is upstreamed/mainline. Triton's bugs have been going over time, but we liked it as a starting point, and
@soumithchintala
-
Unified Software Stack Approach for AI Inference Optimization
By
–
we are truly playing out the "uniform software stack" story. As long as you build a triton backend, and you drive down driver/firmware bugs, that is all that's needed. 99% of the user-level papercuts are taken care of.
It also helps that these are targeted for Inference right -
Triton Backend Enables AMD Code Generation Support
By
–
Triton can generate anything, as long as you write a Triton backend. As of today, upstream Triton can generate AMD code too.
-
Meta’s Chip Strategy: Not a Vendor Business Model
By
–
Meta is not in the chip-selling or chip-renting business
-
Meta Launches MTIAv2 Inference Chip for AI Acceleration
By
–
Meta announces 2nd-gen inference chip MTIAv2.
* 708TF/s Int8 / 353TF/s BF16
* 256MB SRAM, 128GB memory
* 90W TDP. 24 chips per node, 3 nodes per rack.
* standard PyTorch stack (Dynamo, Inductor, Triton) for flexibility Fabbed on TSMC's 5nm process, its fully programmable via the -
MLX and Tinygrad: Open Source ML Frameworks in Complexity Phase
By
–
MLX is nice, early in its complexity phase.
also tinygrad. -
Product Focus: Ethical Comparison Practices and Stakeholder Communication
By
–
focusing on improving your own product, avoid comparisons. if needed to compare, give the other stakeholders a heads-up and a voice.
-
Google DeepMind Leadership Aligns on AI Path Forward
By
–
thanks to @JeffDean and @SingularMattrix for their great leadership today; and @fchollet @dwarak and many others at @GoogleDeepMind for quickly charting a good and aligned path forward together.
We can go back focusing on the unlimited amounts of good work ahead of us.
(Jeff, -
PyTorch Native Benchmarks Vetted and Details Posted
By
–
thanks. this is a response to the "PyTorch native" column. We've vetted the benchmarks and posted details now.
-
FP32 vs TF32 Benchmarking: Compiler Optimization Debate
By
–
This is honestly a baffling response. You cant be saying that benchmarking FP32 vs TF32 (just the dtype) is a "compiler optimization".
Honestly, I'm lost at the face of so much evidence, how you're just still sticking to your story.