> local llms 101 > running a model = inference (using model weights)
> inference = predicting the next token based on your input plus all tokens generated so far
> together, these make up the "sequence" > tokens ≠ words
> they're the chunks representing the text a model sees
>
@theahmadosman
-

Local LLMs 101: Inference, Tokens, and Sequence Explained
By
–
-
Qwen 3.5 4B Runs Locally on iPhone 17 Pro Max with Vision
By
–
Watch Qwen 3.5 4B running locally on my iPhone 17 Pro Max analyze a picture of my old local LLM archive
— Ahmad (@TheAhmadOsman) 13 mars 2026
local AI HAS COMR SO FAR https://t.co/Ic60Fkhyg4 pic.twitter.com/7F46jSnuYMWatch Qwen 3.5 4B running locally on my iPhone 17 Pro Max analyze a picture of my old local LLM archive local AI HAS COMR SO FAR
-
Meta Fell Off: Code Llama Already Forgotten
By
–
i also completely forgot Code Llama was a thing Meta really fell off
-

Google Will Reach AGI First Says Industry Observer
By
–
i am now convinced that Google will get to AGI first
-
GPU Tiers Determine Hiring and AI Agent Productivity
By
–
no, you will be hired based on how many GPUs you have (and which ones as well) higher # of GPUs at higher tiers
means better LLMs
which means better Agentic tooling
and higher productivity -

Major AI Releases Despite Resource Constraints in 100 Days
By
–
crazy how they just keep finding a way to do it with restricted resources look at this list for the major releases in the last ~100 days alone
-
Amazing AMA Session with Open Source Contributors
By
–
This is going to be an amazing AMA Thank you for joining us and for all your opensource efforts!
-

Moonshot AI AMA: Kimi K2 Thinking Model on r/LocalLLaMA
By
–
join us this Monday on r/LocalLLaMA for an AMA with Moonshot AI, the lab behind the SoTA model Kimi K2 Thinking i am genuinely excited for this one, make sure you don't miss it Monday 8am-11am PST