Phone model. Here's why: 1. An o3-mini-level oss model is short-term inevitable. It will happen from somewhere, and it will happen soon. 2. Small local models still have a long way to go. A truly great <8B model will unlock many new uses: edge devices, small projects, etc.
LLMS
-
Elon Musk’s Grok3 Achievement Signals Bigger AI Role
By
–
What @elonmusk achieved with his team to build #Grok3 is a superhuman feat of engineering.
— Nina Schick (@NinaDSchick) 19 février 2025
Could anyone else have pulled it off? Jensen Huang said it was ‘singular’.
That all means that Elon Musk will be a playing an even bigger role in frontier development of AI.
He… pic.twitter.com/tag6eY6De1What @elonmusk achieved with his team to build #Grok3 is a superhuman feat of engineering. Could anyone else have pulled it off? Jensen Huang said it was ‘singular’. That all means that Elon Musk will be a playing an even bigger role in frontier development of AI. He
-
LangChain CEO Event Atlanta February 27th
By
–
LangChain in Atlanta! Join us next Thursday, February 27th for an evening of AI at the Honeywell offices with LangChain CEO, @hwchase17
-

Ultra-Scale Playbook: 5D Parallelism and CUDA Optimization Guide
By
–
After 6+ months in the making and burning over a year of GPU compute time, we're super excited to finally release the "Ultra-Scale Playbook" Check it out here: http://
hf.co/spaces/nanotro
n/ultrascale-playbook
… A free, open-source, book to learn everything about 5D parallelism, ZeRO, fast CUDA kernels, -

Phi-4 Model Release: 40% Synthetic Data Composition
By
–
Phi-4! (
https://
arxiv.org/abs/2412.08905). 40% synthetic. -

DeepSeek-R1 Game-Changer for Enterprise AI Deployment
By
–
In his interview with @CliffSaran at @ComputerWeekly
, @RodrigoLiang explained why DeepSeek-R1 is a game-changer for enterprises aiming to deploy cutting-edge AI without the budget strain. Read the full story here: https://
computerweekly.com/news/366619398
/DeepSeek-R1-Budgeting-challenges-for-on-premise-deployments
… #AI -

SFT and RL Training Pipeline: Limitations of Cold Start Approach
By
–
Yes & no. They had SFT (cold start) → RL → SFT (CoT + knowledge) → RL. Not sure if they focused on providing solutions to hard problems via SFT CoT though, because that SFT data was just generated from the previous model (so if that model doesn't solve it, then you are stuck)
-
Alternating SFT and RL Training for Improved Model Policy
By
–
I wonder if alternating SFT (with correct solutions) and RL can help here. I.e, SFT → RL → SFT → RL. SFT would help with improving the initial policy and help making the exploration less random perhaps, plus it expands/refines the search space?
-
Model Performance Comparisons: Size Normalization and Inference Scaling Challenges
By
–
Model performance comparisons are tricky without knowing exact sizes; it could be an apples-to-oranges case. Plus, you can always push numbers with inference scaling. Normalizing by tokens/sec would be more meaningful.
-

Grok 3 context window size clarification
By
–
What is Grok 3 context window again? The post from @arnogau referring to a 1M context window has been removed
