Long uninterrupted runs are the new benchmark. A model that doesn't stop and start on a 50k-token refactor is shipping a different product than one that does, even if they score similar on short tasks. 4.6 stopping was masking how sensitive the previous loop was to noise.
RESEARCH
-
Adaptive Thinking Trade-off: Token Burn vs Performance Regression
By
–
The adaptive thinking burns more tokens and the results drop. That's a regression no matter how the marketing reads. The real question is whether this is a calibration bug fixable in a point patch or a deeper reward-shaping choice that won't roll back… in any case, not so happy
-
API Access Restrictions Shape Open vs Closed Research Agendas
By
–
Frontier API access shrinking is already shaping research agendas. The labs that publish and open-weight find the bugs the closed ones can't, because no single research setup is broad enough.
-
Deep Stack Attention Hijacking: Beyond Prompt Injection Vulnerabilities
By
–
a lot of injection research focuses on the prompt surface and misses that the vulnerability is actually in how routing attention gets hijacked deep in the stack
-
Rosalind Project: Ambitious Drug Discovery Effort and Research Implications
By
–
Thanks Maggie for reaching out. Rosalind seems like a very ambitious and worthwhile effort, and I am very hopeful about what may come out of it. However, drug discovery is a far cry from deep fundamental science research. Also, based on my first hand experience with some of the
-
Rosalind Project: Drug Discovery and Fundamental AI Research Challenges
By
–
Thanks Maggie for reaching out. Rosalind seems like a very ambitious and worthwhile effort, and I am very hopeful about what may come out of it. However, drug discovery is a far cry from deep fundamental science research. Also, based on my first hand experience with some of the
-
Opus 4.7 Shows Improvement But Still Underperforms 4.6
By
–
Opus 4.7 does seem to have improved, and its adaptive thinking now uses more tokens. However, compared to Opus 4.6, it still performs significantly worse.
-

Pop-up Attacks Hijack AI Assistants: LaSM Defense Method
By
–
What if a simple pop-up could completely hijack your AI assistant? Researchers from Shanghai Jiao Tong University have a new fix. They found that pop-up attacks misdirect an AI's "attention" in specific layers of its neural network. Their method, LaSM, defends by
-
Paradigm Shifts and Resistance to Technological Change
By
–
Yup. Every time in my life there is a paradigm shift droves of resistors come out. More this time.