huh, that needs extra work for Claude? Codex just does that by default.
LLMS
-

Who designs RL agent’s training environment: practitioner or policy?
By
–
Who should design the training environment for an RL agent, the practitioner or the policy itself? RL pipelines for LLMs usually rely on manually redesigned environments between stages, with practitioners guessing which configuration will best improve the current policy. This
-

Stem: Efficient Long-Context LLM with Token Position-Decay
By
–
What if your LLM could process long contexts without the quadratic attention bottleneck? Enter Stem: a plug-and-play sparsity module that rethinks causal information flow. It uses a token position‑decay strategy (keeping early tokens for recursive dependencies) and an
-
GLM 5.2: Second Most Important Open-Source AI Moment After Qwen 3.5
By
–
GLM 5.2 was the 2nd most important moment in Opensource AI after Qwen 3.5 27B in 2026
-

Autonomous Agents and Task Repetition
By
–
Outstanding paper on computer-using agents. (bookmark it) Computer-using agents drive real software through the screen, but they solve every task from scratch. Ask one to repeat a task, and it re-reads the screen and re-reasons every tap, paying the full cost again. PreAct
-
Codex excels at creating and manipulating Excel and Google Sheets
By
–
Codex is remarkably well at creating and manipulating sheets in excel, google sheets! I’ve had it create complex pivots, dashboards and use sheets as a way to keep track of tasks Works remarkably well, best part is that you don’t need a specific skill for this, GPT 5.5 is
-

Trump imposes strict conditions for the re-release of Fable 5
By
–
This seems very bad for a future re-release of Fable 5. "Trump administration officials told WIRED that if Anthropic wants to re-release Fable 5, it must ensure that the model's guardrails cannot be bypassed. Security experts
-
OSS eventually good enough for 99% consumer tasks, not highest-value
By
–
Never say never, but I'm highly skeptical. OSS *will* eventually be good enough for 99% of consumer tasks, but the highest-value tasks will require the most powerful models.
-
You can vibe code as long as you test after
By
–
Good point and yes partly it's that and people vibe coding But you can also just ask AI to add more test so things keep working well You can vibe code all you want as long as you test after
-
China and its Mythos model: speed is crucial
By
–
It is certainly not a coincidence, I think the same thing. The question, in the end, is how fast China will release its own Mythos-class model. Almost everything depends on it.