Ooof. I don't know if 128K is possible on the 70B, unless someone can get me a rig with like 16 or more A/H100s. 32K is likely doable though.
AI HARDWARE
-
Seeking $500 Compute Sponsorship for AI Project
By
–
If someone wants to sponsor the compute, I can do it! Would only need like $500 of compute.
-
Running AI Applications on Budget Laptops: RAM Requirements
By
–
Presumably this one but I'm surprised it runs at all on a cheap laptop, @damnkittyworks how much RAM do you have?
-
Groq’s Time to First Token Performance Needs Improvement
By
–
If Groq gets their time to first token down a bit, probably not. They still take maybe a half second to a second to get you that first token.
-

Nvidia plummets: TL20 index records worst day ever
By
–
The TL Podcast for April 21st: Nvidia takes a tumble Nvidia fell on Friday by the most in the two and a half years of the TL20, giving the group its worst day ever. In the absence of any particular news, this event was a healthy right-sizing. Nvidia still has no real competition
-
Energy-Efficient AI Inference Performance at Hannover Messe
By
–
Excited to be at @hannover_messe with @NGen_Canada
! Visit us in Hall 17, Stand D42 to see how we can maximize your #AI inference performance while still being kind to the planet. #EarthDay Chat with our experts & witness the future of energy-centric AI firsthand! #HM24 -
CUDA CuDNN Python Startup Costs and Compilation Overhead
By
–
its unclear to me even for code because compile-time and startup costs (even for a CUDA/CuDNN loading Python program) is 1 second.
If you strip out the binary blobs that need to load (say business logic filled program), then its quicker, but i'm not sure if there aren't other -

Competition Against Nvidia in AI Chip Market Intensifies
By
–
My good friend Eric Savitz has published a lovely feature on the competition against Nvidia for Barron's, with excellent data and examples. I've been writing about the AI chip competition to Nvidia for almost a decade. In all that time, no one has put a dent in Nvidia's
-

Meta-Llama-3-70B runs on M2 Mac with 64GB RAM
By
–
OK, can confirm that it works on an M2 with 64GB of RAM I had to quit Firefox and VS Code to free up enough RAM for it to run – Meta-Llama-3-70B-Instruct.Q4_0.llamafile now uses 38GB of RAM and runs at about 7.5 tokens a second