GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repeated misuse, then spent weeks hardening the system with human red teaming and over 700,000 A100-equivalent GPU hours of automated testing.
LLMS
-

GPT-5.6 Sol sets new state-of-the-art on Terminal-Bench 2.1
By
–
GPT‑5.6 Sol sets a new state of the art on Terminal‑Bench 2.1, which tests complex command-line workflows requiring planning, iteration, and tool coordination.
-
Static benchmarks measure memorization, not intelligence
By
–
If your benchmark relies on a static dataset or sampling from a static distribution densely known at training time, then it is fundamentally measuring memorization/retrieval. Which might be fine if you're looking for a retrieval benchmark! But don't confuse it with intelligence.
-

OpenAI announces GPT-5.6 Sol with restricted US access
By
–
BREAKING: OpenAI announced GPT-5.6 Sol! As of today, by U.S. government directive, access is limited to only ~20 pre-approved companies and @every is not on the list. This appears to be a temporary situation while the government races to figure out a long-term policy for
-
Request for Codex session IDs with poor performance
By
–
hey hey morgan – can you send me one of your session IDs from your past codex sessions from today/ yesterday where you felt codex was off?
-

Stanford AI adds realistic personalities to train crisis workers
By
–
Stanford scholars developed a new way of adding human differences back into AI-generated text. With more realistic personalities, AI can simulate patients with specific symptom profiles for training crisis-line workers and clinicians. https://
hai.stanford.edu/news/todays-ai
-talks-like-nobody-new-research-gives-it-real-personality
… -

Open-source Ornith-1.0 models rival frontier agentic coding performance
By
–

Open-source models are rapidly closing the gap on frontier agentic coding. The launch of the Ornith-1.0 family proves you don't need a closed API to get top-tier coding agent performance! Built on top of Gemma 4 and Qwen 3.5, Ornith uses a self-scaffolding RL pipeline to
-

PACE: Training optimizers for the averaged model you return
By
–
"Training for the Model You Return" Most LM pipelines return an EMA or averaged checkpoint, but optimizers still train like the final iterate is what matters. So if the output is an averaged model, can we shape training so that average gets better? This paper introduces PACE,
-
Billions spent on AI, but memory was the real bottleneck
By
–
Do you understand the irony of what just happened?
— Robert Scoble (@Scobleizer) 26 juin 2026
We spent two years and BILLIONS of dollars making AI models smarter. What was actually holding them back was that they forget everything after each session.
It was memory. https://t.co/0krKlMh469Do you understand the irony of what just happened? We spent two years and BILLIONS of dollars making AI models smarter. What was actually holding them back was that they forget everything after each session. It was memory.
-
Anthropic advances study of Claude’s economic impact with hourly data
By
–
To keep pace with AI progress, we're advancing how we study Claude's economic impact. Hourly sampling and survey data show us how the cadences of life shape usage, what people produce with Claude, and how perceptions of AI's impact may be changing.