wdym? we released a new instant model, 5.2 Deep Research, hosted shell, skills and compaction in the API – all of this in just the last 5 days! there’s never been a better time to build w/ OpenAI
LLMS
-
DSA: Learning-Based Token Selection Beyond Fixed Window Attention
By
–
Also, DSA is basically a smarter version of sliding window attention where you "learn" which past tokens to select versus forcing it to be in a specific window
-
DeepSeek V3.2 Success: Why Not Apply This Approach
By
–
I mean why not? It worked pretty well in DeepSeek V3.2
-
Question about sliding window attention implementation
By
–
Hmm I don’t think they use sliding window attn ? Where did you see that?
-
Lean and Natural Language Reasoning in AI Exploration
By
–
Yes to both 🙂 I think Lean is very important, especially in the future. For now, natural language reasoning helps us explore more domains.
-
Expert Size and Number Match DeepSeek V3 Specifications
By
–
Probably also should have added that
– the expert size (2048) and number (256) is now exactly the same as in DeepSeek V3 / V3.2. -
Vibe coding ML library refactor fails with no speedup
By
–
Yesterday I tried to vibe code a refactor of an ML library into a new more efficient framework. I followed some of the "best practices" I heard of here: created the refactor plan with Claude and then implemented it in Codex. It was a disaster. No meaningful speedup for any of
-

GLM-5 Architecture: Weights Released, Key Features Unveiled
By
–
The weights are out! Here's the GLM-5 architecture comparison. GLM-5 is: – bigger than its predecessor (mainly more experts) but has rel. similar active parameter counts – uses multi-head latent attention – uses DeepSeek Sparse Attention
-
Musk’s manipulation of generative AI on Grok
By
–
Yes, of course, generative AI is entirely manipulable. Musk has demonstrated this several times by having algorithms modified to change Grok's editorial line on certain subjects in just a few hours.
