Fuck it – it’s raining smol LMs – SmolLM2 1.7B – beats Qwen 2.5 1.5B & Llama 3.21B, Apache 2.0 licensed, trained on 11 Trillion tokens > 135M, 360M, 1.7B parameter model
> Trained on FineWeb-Edu, DCLM, The Stack, along w/ new mathematics and coding datasets
> Specialises in
@reach_vb
-

SmolLM2 1.7B Beats Larger Models with Apache 2.0 License
By
–
-
VRAM Requirements for AI Models Across Hardware Architectures
By
–
It should work on CPU/ CUDA/ MPS across backends, w.r.t hardware requirements: 1B should take roughly 2GB VRAM to load in fp16/ bf16.
600M should take 1.2 GB VRAM
350M – ~700MB VRAM
125 – ~250MB VRAM Ofcourse at lower quants Q4/ Q8 you reduce this even further. -

Tiny AI Models Running on Everyday Devices
By
–
Models so smol that thay'd even run on your toaster!
-

Meta Releases MobileLLM: Depth Critical Over Width
By
–
Meta released MobileLLM – 125M, 350M, 600M, 1B model checkpoints! Notes on the release: Depth vs. Width: Contrary to the scaling law (Kaplan et al., 2020), depth is more critical than width for small LLMs, enhancing abstract concept capture and final performance Embedding
-
Hugging Face Releases AutoTrain Advanced Open Source Code
By
–
And the code lives free for anyone to use too! https://
github.com/huggingface/au
totrain-advanced
… -
The Power and Possibilities of Open Source Software
By
–
The beauty of open source is that you can do all of that and more.
-
Community Praise for Open Source and Scientific Contributions
By
–
Class apart! – I love all that y’all do for open source & science!
