8/ This too will be one of the great infrastructure projects of our time—ensuring we build the means of production of new data across the frontier to fuel AI progress.
AI HARDWARE
-
Scaling Compute and Data: The Infrastructure Challenge Ahead
By
–
5/ The next 4 years, like the last 4 years, will be all about scaling. Much of progress depends on our ability to keep scaling compute + data exponentially. This is no easy feat. Scaling data + compute will be some of the largest infrastructure projects of our time.
-

Cerebras CEO Breaks GPU Barriers for AI Model Training
By
–
Check out our CEO Andrew Feldman’s keynote on how Cerebras has broken through GPU barriers and makes advanced AI model training dead simple at the Mint Digital Innovation Summit 2024. Watch Andrew’s talk: https://
livemint.com/industry/ai-is
-easy-to-describe-challenging-to-do-cerebras-systems-ceo-andrew-feldman-11716561044183.html
… Contact Cerebras to accelerate your AI -
H100 and FP8 optimizations for performance improvement
By
–
Nice! H100 is a great "free win" to bring this down.
Turning on fp8 for GEMMs would be the other source of really solid improvement, imminently -
Intel AI achieves outstanding performance milestone
By
–
That was one of the best of all time @IntelAI
-
AI Revolution Began with ENIAC Computer in 1945
By
–
The AI revolution started with the ENIAC in 1945.
-
Dealing with random MPI hangs in large model training
By
–
I thought I didn't have to deal with these, but already the 350M model (14 hours of 8 GPUs working) sometimes randomly hangs with a cryptic MPI error once in a while. So I have to put the whole optimization into a `while 1` loop and a script that watches the log file and sends
-
Free Meetup: Building LLM Serving Platforms at Scale with NVIDIA
By
–
Bay Area Friends: Join us tomorrow at Orchestrating #GenAI Apps Meetup with NVIDIA and NetApp. We'll be discussing what it takes to build an #LLM serving platform at scale. Best of all it's free to attend, save your spot:
-
Single Node Training vs Large-Scale Multi-GPU Distributed Runs
By
–
But those were also much much bigger runs, so it's a lot more impressive. This was on a single node so you don't need to deal with any cross-node interconnect. It starts to get a lot more fun when you have to keep track of O(10,000) GPUs all at once. For a very specific
-

Nvidia Valuation Analysis: Can Stock Rally Higher?
By
–
No, Nvidia is not overpriced. The TL Podcast for May 27th: Sizing up Nvidia Nvidia had a fantastic quarter, does its valuation allow the stock to go higher? And Snowflake’s CFO explains the complexities of its accounting. https://
thetechnologyletter.com/the-posts/the-
tl-podcast-for-may-27th-sizing-up-nvidia
… $NVDA $SNOW