We had a great time at MHacks this past weekend. Kudos to the WiLi Watch Team – Alexandra Enders, Anshul Mohanty, Austyn Nguyen, Ethan McKean – who won the Best on Groq prize.
Check out their project here: https://
hubs.ly/Q02RXnLN0
AI HARDWARE
-

Groq Prize Winners Announced at MHacks Hackathon Event
By
–
-

Untether AI Partners Ampere for Energy-Efficient AI Inference
By
–
We've joined the AI Platform Alliance with @AmpereComputing
! See our partnership in action at #Yotta2024 – we're showcasing our record-breaking speedAI240 Slim PCIe card. Visit us in #601F, Oct 7-9, to see the future of sustainable, high-performance #AI. https://
untether.ai/untether-ai-jo
ins-forces-with-ampere-to-provide-arm-based-energy-efficient-ai-inference-solutions/
… -
RAS Blackwell GPU Isolation and Host Boot Issues
By
–
RAS is blackwell+, so TBD. no single GPU isolation, whole host gets booted.
-
Training Machine Learning Models on 10k H100 GPUs
By
–
"How to train a model on 10k H100 GPUs?"
has now been immortalized on my blog: https://
soumith.ch/blog/2024-10-0
2-training-10k-scale.md.html
… -
HBM Memory in Switches: Critical Infrastructure for AI Scale
By
–
one more thing to add.
at this scale, we also have to adjust the actual packet routing algorithms in our switches and NICs, to be able to load-balance well. Did you know switches have to have significant HBM memory as well (not just GPUs) because as the packets queue up, they -
Training Models at Scale: 10K+ H100 Infrastructure Guide
By
–
quick mini-post I wrote for @francoisfleuret broadly summarizing the things one needs to do to train models on 10k+ H100s.
-
Scaling AI Training Across Thousands of H100 GPUs
By
–
There's three parts. 1. Fitting as large of a network and as large of a batch-size as possible onto the 10k/100k/1m H100s — parallelizing and using memory-saving tricks.
2. Communicating state between these GPUs as quickly as possible
3. Recovering from failures (hardware, -
AlphaChip: How AI Redefines Chip Design
By
–
[#Article] AlphaChip: How DeepMind's AI Redefines Chip Design Standards https://actuia.com/actualite/alphachip-comment-lia-de-deepmind-redefinit-les-normes-de-la-conception-des-puces/
… #AI #ArtificialIntelligence -
AI Productivity Gains: From Consultants to Quantum Computing
By
–
Les sources du FT sont ici : https://
computerworld.com/article/354342
1/ai-could-be-a-good-replacement-for-bad-bosses-and-expensive-consultants.html
… https://
next.ink/brief_article/
ibm-ouvre-son-premier-datacenter-quantique-en-europe/
… https://
bloomberg.com/news/articles/
2024-10-01/fed-s-cook-sees-ai-boosting-productivity-how-much-still-unclear
… https://
techcrunch.com/2024/10/01/mic
rosoft-copilot-can-now-read-your-screen-think-deeper-and-speak-aloud-to-you/
… https://
zdnet.fr/actualites/la-
startupeuse-julie-huguet-nouveau-visage-de-la-french-tech-398673.htm#xtor=RSS-1
…