GPT-5.5 cranking out 30k lines of QML for the Omarchy 4 branch + nailing subtle agentic reasoning!!
LLMS
-

Claude MD File: Top GitHub Repo Uses Karpathy’s 4 LLM Principles for AI Control
By
–
DID YOU KNOW THE #1 GITHUB TRENDING REPO (146K+ STARS) IS LITERALLY JUST A CLAUDE MD FILE? It uses @karpathy
's 4 LLM principles to keep AI in check: → Seek clarity: Always ask before making assumptions.
→ Stay minimal: Avoid bloat and keep things simple.
→ Edit -
Thousands of tokens per second across parallel requests
By
–
1000s of toks/sec across a dozen parallel requests if not more
-
Gemini 3.5 Flash: Eager Model Built for Real-World Usefulness
By
–
Gemini 3.5 Flash is definitely an over eager model, a bit of over correction from the “Gemini laziness” feedback, but definitely built to be real world useful!
-
LLM trained on whole Internet is inefficient for language understanding
By
–
Learning an LLM from the whole Internet is a spectacularly inefficient way to understand language.
-
Critique of local AI performance metrics as TikTok-style grifting
By
–
No, I am saying that this is the equivalent of TikTok brain rot for Local AI Longest prompt was 368 tons, longest output was 3k and some grifters out there are saying that this is great results of 45tok/s lol If you cannot see how this is performative slop / grifting then
-
Sarcastic commentary on short prompt and costly hardware
By
–
I mean, dude's longest prompt was 368 tokens and longest output was 4k He's literally crying next to that hardware and hoping to get enough engagement that elonbux can pay for his next month installment on those desk warmers
-
GPT 5.5 improving while Claude models worsening, no clear winner
By
–
GPT 5.5 seems to be improving in that direction now, and Claude models are getting worse at it, so I don't think there's a clear winner now.
-
Gemini Flash 3.5 criticized for prioritizing evals over user helpfulness
By
–
Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's been trained to max evals, not to be helpful to humans. It goes off and does random crap "for me" rather than just doing what I asked.
-
Need better evaluation of models for human-AI cooperation
By
–
We desperately need better ways of evaluating models. Something that shows how helpful they are at working hand-in-hand with humans to help them get stuff done in a cooperative/iterative way. The Claude models have consistently been better at this, and the market rewards that.