+1 to same compiler flow on our chips as mainline PyTorch. At Meta, we don't have a separate copy of PyTorch or Dynamo/Inductor. The source of truth is Github and everything is upstreamed/mainline. Triton's bugs have been going over time, but we liked it as a starting point, and
AI HARDWARE
-
Faust’s Infernal Bargain: Tech Giants and Environmental Trade-offs
By
–
While I admire her spirit, Faust would have made a deal with a16z to build a data center in place of the redwood tree
-
Unified Software Stack Approach for AI Inference Optimization
By
–
we are truly playing out the "uniform software stack" story. As long as you build a triton backend, and you drive down driver/firmware bugs, that is all that's needed. 99% of the user-level papercuts are taken care of.
It also helps that these are targeted for Inference right -
Triton Backend Enables AMD Code Generation Support
By
–
Triton can generate anything, as long as you write a Triton backend. As of today, upstream Triton can generate AMD code too.
-
Meta’s Chip Strategy: Not a Vendor Business Model
By
–
Meta is not in the chip-selling or chip-renting business
-

NVIDIA Stock Correction: GPU Demand Concerns Amid AI Model Shifts
By
–
NVIDIA stock briefly hit “correction" territory. The GPU maker's stock fell 10% from ATH at Tuesday's close. It's still up 200% year-over-year, though. DA Davidson fears a "shrinking" of AI models in future will lower demand for Nvidia's powerful chips. https://
cnbc.com/2024/04/10/nvi
dia-nvda-stock-down-10percent-from-highs-in-correction-territory.html
… -
Simplified Management of Serverless and Dedicated LLM Deployments
By
–
ICYI: Managing #serverelss and dedicated #LLM deployments has never been easier! 🚀
— Predibase by Rubrik (@predibase) 10 avril 2024
😎 see all #deployments and status in one place
🆕 create new dedicated deployments with just a few clicks
🛠️ select the right #GPU for the job and customize autoscalinghttps://t.co/W9hCTKaySr pic.twitter.com/zImeI45nNFICYI: Managing #serverelss and dedicated #LLM deployments has never been easier! see all #deployments and status in one place create new dedicated deployments with just a few clicks select the right #GPU for the job and customize autoscaling https://
pbase.ai/449UUXT -
Meta Launches MTIAv2 Inference Chip for AI Acceleration
By
–
Meta announces 2nd-gen inference chip MTIAv2.
* 708TF/s Int8 / 353TF/s BF16
* 256MB SRAM, 128GB memory
* 90W TDP. 24 chips per node, 3 nodes per rack.
* standard PyTorch stack (Dynamo, Inductor, Triton) for flexibility Fabbed on TSMC's 5nm process, its fully programmable via the -

Meta Llama 3 Launch, AI Chips, and New Tools
By
–
Top stories in AI today: -Meta confirms Llama 3 coming within the month
-AI chip wars heat up with Intel Gaudi 3
-Make presentations 10x faster with AI
-Cohere’s Command R+ climbs leaderboard
-6 new AI tools & 4 new AI jobs Read more: http://
therundown.ai/p/metas-llama-
3-waiting-game
… -

AI Breakthroughs in Sports, Edge Computing, VRAN and 5G at MWC
By
–
Webcast ft. Chi Tran, John Yates, Steve Dinkins & Kevin Huisuk Hong: AI Breakthroughs at MWC for #Sports, #Edge, VRAN & 5G
by @Ronald_vanLoon | Register To Get Notified: https://
bit.ly/3U0rK9S #IntelAmbassador @Intel @IntelBusiness @IntelEdge #MWC24 #5G #Networking #IoT Cc: