> llama-like architecture
> dense transformer > rotary only (no positional embeddings) > qk norm > untied embedding/unembedding > norm after token embedding > relu² mlp > no biases in linears > no learnable rmsnorm params > mqa > logit softcap > optimizer =
@theahmadosman
-
Advanced Llama Architecture: Rotary Embeddings and ReLU² MLP
By
–
-
Karpathy’s Fundamental AI Resource: Essential Guide for Beginners
By
–
this is an amazing gift to everyone who's just starting straight to the point and has the fundamentals covered in the right order (no surprise though since it's karpathy)
-

Karpathy releases new LLM training project under 8000 lines
By
–
> karpathy just released a new project > in it, and in under 8000 lines of code you get to: > train the tokenizer using a new rust implementation
> pretrain a transformer llm on fineweb, evaluate core score across a number of metrics > midtrain on user-assistant -
From Zero to Attention Mechanisms: Demystifying LLM Knowledge
By
–
– you are
– a random CS grad with 0 clue how LLMs work
– get tired of people gatekeeping with big words and tiny GPUs
– decide to go full monk mode
– 2 years later i can explain attention mechanisms at parties and ruin them – here’s the forbidden knowledge map
– top to bottom, -
Invest in Yourself: The Ultimate Betting Strategy for Success
By
–
betting on this, betting on that but when was the last time you tried betting on yourself, anon?
-
Qwen 3 release questioned technical report limitations
By
–
I think the only one I might question the significance of in your list is Qwen 3. They didn't even release a base model. The technical report was quite disappointing as well.
-
vLLM Limitations in Broader AI Landscape
By
–
vLLM is a small part of the whole picture. They're massively lacking.
-
Inference Limited Role in AI Development Picture
By
–
Inference is a small part of the picture. Not comparable.
-
GPU Market Viability: CUDA Support and Hardware Trends
By
–
No. And people are still arguing with me lol. You still need to be aware of what's happening in CUDA, when a certain GPU support is expected to drop, etc. But It's a joke if you think GPUs are a bubble.
-
GPU Bull Case Remains Strong for AI Infrastructure
By
–
It's still a bull-case for GPUs in my opinion though