It should work on CPU/ CUDA/ MPS across backends, w.r.t hardware requirements: 1B should take roughly 2GB VRAM to load in fp16/ bf16.
600M should take 1.2 GB VRAM
350M – ~700MB VRAM
125 – ~250MB VRAM Ofcourse at lower quants Q4/ Q8 you reduce this even further.
LLMS
-
VRAM Requirements for AI Models Across Hardware Architectures
By
–
-
A user pays for multiple AI services and admits being ‘crazy’
By
–
I pay:
Gemini for business
Claude
ChatGPT
API OpenAI
API Claude
API Mistral
API Gemini
Cursor + supplement
Copilot Pro Don't worry, when you think you're crazy, there's always more crazy…. -

Tiny AI Models Running on Everyday Devices
By
–
Models so smol that thay'd even run on your toaster!
-

Meta Releases MobileLLM: Depth Critical Over Width
By
–
Meta released MobileLLM – 125M, 350M, 600M, 1B model checkpoints! Notes on the release: Depth vs. Width: Contrary to the scaling law (Kaplan et al., 2020), depth is more critical than width for small LLMs, enhancing abstract concept capture and final performance Embedding
-
SimpleQA Evaluation Tool Update on GitHub
By
–
https://
github.com/openai/simple-
evals/blob/main/simpleqa_eval.py#L104C19-L104C97
… Updated here! -

Building Large-Scale Clusters to Train and Release Llama
By
–
its super fun to build very very large clusters, train llama on them, and release it for y'all to enjoy — and talk in great detail about how we did it!
It's also really fun to partner with @Ahmad_Al_Dahle in creating this disruptive chaos Join us, there's lots of work to do! -
Non-Autoregressive TTS Model Eliminates Need for Alignment
By
–
a fully non-autoregressive TTS model that eliminates the need for explicit alignment information between text and speech supervision, as well as phone-level duration prediction
-

AI and Machine Learning Impact Analysis 2024
By
–
blog/2024.09.impact.md at main · okhat/blog https://
bit.ly/3ZJQ8jl
#AI #MachineLearning #DeepLearning #LLMs #DataScience
