LLM sceptics have predicted the last 7 of 0 walls
LLMS
-

Cerebras Wafer Scale Advantage Over NVIDIA Groq Inference Chips
By
–
Problem solved. ✅ Andrew Feldman (@andrewdfeldman) NVIDIA's biggest GTC announcement was a $20 billion bet on the same problem we solved 6 years ago. Their next-gen inference chip – not available yet – has 140x less memory bandwidth than @cerebras. To run a single 2 trillion parameter model, you need 2,000+ Groq chips. On Cerebras, that's just over 20 wafers. Even paired with GPUs, Groq maxes out at ~1,000 tokens per second. We run at thousands of tokens per second today. And every day. In production now. Why? When you connect 2,000 chips together, every interconnect has latency. Every cable has overhead. It doesn't matter what your memory bandwidth is on paper if you're bottlenecked by the wiring between thousands of tiny chips. We solved this with wafer scale. One integrated system. Little interconnect tax. Jensen told the world that fast inference is where the value is. He’s right – it’s why the world’s leading AI companies and hyperscalers are choosing Cerebras. — https://nitter.net/andrewdfeldman/status/2034015373595672594#m
-

50 ML Projects to Understand LLMs and Transformer Mechanisms
By
–
50 ML projects to understand LLMs — Investigate transformer mechanisms through data analysis, visualization, and experimentation: http://
amzn.to/4aPfP7q
—————
#AI #GenAI #MachineLearning #DataScientist #DataScience -
Frontier AI Coding API Pricing at $0.50 per Million Tokens
By
–
$0.50/M input for frontier-level coding is next level for the cost side. Curious how it stacks up for longer multi-step tasks!
-
Continued Pretraining and Scaled RL for Specialized AI Models
By
–
Continued pretraining + scaled RL is a combo I keep seeing more of. The niche specialization angle is underrated!
-

AMD CEO Lisa Su meets Upstage in Seoul for AI partnership
By
–

Honored to meet AMD CEO @LisaSu in Seoul. We're adopting AMD's Instinct MI355X GPU to power our Upstage Solar LLM and Korea's sovereign AI. The future of AI is being built right here. 🇰🇷🚀 #AMD #AI #Upstage #SolarLLM
-

Beyond VRAM: Understanding LLM Inference Bottlenecks Locally
By
–
slap an api on this for our agents
-
Essay Anticipates Security Nightmare from LLMs and Coding Agents
By
–
and see the essay that anticipated the chaos: https://
open.substack.com/pub/garymarcus
/p/llms-coding-agents-security-nightmare?utm_campaign=post-expanded-share&utm_medium=web
… -
LLMs, Coding Agents, and Security Nightmare
By
–
https://
garymarcus.substack.com/p/llms-coding-
agents-security-nightmare
… -

Parallel-Probe: Faster AI Reasoning Without Performance Loss
By
–
Can we make AI reasoning much faster and more efficient without losing its smarts? Researchers from the University of Maryland, Washington University in St. Louis, and UNC Chapel Hill introduce Parallel-Probe. This innovative, training-free controller uses "2D probing" to