> E.g. Llama 3 405B used 30.8M GPU-hours, while DeepSeek-V3 looks to be a stronger model at only 2.8M GPU-hours (~11X less compute). Super interesting! And DeepSeek was trained in H800’s which are probably also a tad (or noticeably?) slower than Meta’s H100’s.
HARDWARE
-
Nvidia B200 availability reflects actual market demand
By
–
If you really wanted a Nvidia B200 you would already have one
-

Frontier Model Development Cost Drops to 5.5 Million USD
By
–
Scarcity breeds Innovation – cost to build a frontier model – 5.5 Million USD In a way, it's the maximum it'd be (Note: H800s have ~2x slower chip-to-chip data transfer) This cost, will only go down further and further as we continue to find newer walls to scale!
-
Qwen and Meta GPU Mobilization in AI Competition
By
–
I wouldn't discount Qwen or Meta either – at least the latter is mobilising metric fk ton of GPUs
-
GPU Deployment Challenges for AI Model Infrastructure
By
–
Not sure if we have the spare GPUs atm, doesn’t look as likely in the short term tho – since it might need some updates in TGI. Folks from @hyperbolic_labs mentioned (on shared slack) they might deploy it!
-

AI Hackathons: Google Cloud MLB and Windows Snapdragon Opportunities
By
–
#AI Hackathons hosted by @devpost … Google Cloud x MLB(TM) Hackathon – Building with Gemini Models PRIZES: $98,700 DEADLINE: Feb 4, 2025 Build the future of baseball using Google Cloud.
JOIN THE HACKATHON:
https://
bit.ly/googlecloudaic
hallengei
… Windows on Snapdragon AI Hackathon -
Device Missing USB-C Cable Standard Accessory
By
–
but it doesn’t ship with a usb-c charging cable, which is pretty dumb
-
Meta Ray-Bans Launch Receives Positive User Feedback
By
–
meta raybans first reactions “It’s great” “It’s amazing”
-
GPU upgrade wish for SE2 model improvement
By
–
Thanks! Wish you a new GPU, SE2 needs it and is worth it
