To preserve chain-of-thought (CoT) monitorability, we must be able to measure it. We built a framework + evaluation suite to measure CoT monitorability — 13 evaluations across 24 environments — so that we can actually tell when models verbalize targeted aspects of their
LLMS
-
OpenAI Model Spec: Intended Behavior for AI Models
By
–
The Model Spec — intended behavior for the models that power OpenAI’s products:
-

NotebookLM adds AI-powered Data Table extraction and export features
By
–

Data Table artefact is rolling out on NotebookLM, where it can generate a structured table with your data. This artefact can be exported into Google Sheets as well. Notes and Reports can be exported into Docs and Sheets, too.
-
Testing world model capabilities in language models like GPT
By
–
Maybe we should come up with some more world model tests that could show it more definitively, one thing that is hard is whether we are just testing the language component of the model which confuses things. I haven't obviously noticed that GPT was worse at world model stuff, but
-

Claude Code to Introduce Shared Session Functionality
By
–
Users will be able to share Claude Code sessions in the future. This feature will potentially allow users to review and continue shared conversations. Vibe Code review
-
GPT-5.2-Codex Advances Agentic Coding and Cybersecurity Standards
By
–
GPT-5.2-Codex is now available in Codex. It sets a new standard for agentic coding in real-world software development and defensive cybersecurity. It also delivers more reliable performance on complex tasks and scales effectively across large projects.
-

Codex Advances in Security Vulnerability Detection and Defensive Programs
By
–
Codex also getting very good at finding security vulnerabilities. We're exploring trusted access programs for defensive cybersecurity work, opening up the opportunity for enterprises and the open-source community to produce more secure code. More here: https://
openai.com/index/introduc
ing-gpt-5-2-codex/
… -

GPT-5.2-Codex: Advanced Model for Long-Horizon Agentic Coding
By
–
Just launched GPT-5.2-Codex! The best model for long-horizon agentic coding, including strong performance on refactors and migrations. Codex becoming very magical.
-
Claude Chrome Extension Gains DOM and Network Inspection Capabilities
By
–
Claude in Chrome extension is now available to all paid plans, including Claude Pro users.
— 🚨 AI News | TestingCatalog (@testingcatalog) 18 décembre 2025
Claude in Chrome can access the page DOM and inspect network requests, too.
Exactly what I need 👀 https://t.co/iaOAkPAZxY pic.twitter.com/IPfpH9FT5tClaude in Chrome extension is now available to all paid plans, including Claude Pro users. Claude in Chrome can access the page DOM and inspect network requests, too. Exactly what I need
-
Gemini Progress Recap with Oriol Vinyals Jeff Dean and Noam Shazeer
By
–
Recapping an incredibly year of Gemini progress with @OriolVinyalsML @JeffDean and @NoamShazeer live, join us : ) nitter.net/i/spaces/1eaJbjvBOooJX
→ View original post on X — @oriolvinyalsml, 2025-12-18 19:53 UTC