For this research, we analyzed only ChatGPT conversations from users who allow their data to be used to improve models. Before analysis, we removed account-linked identifiers and identifiable information, and we report only aggregate findings.
RESEARCH
-

OpenAI analyzes opt-in ChatGPT conversations after anonymization
By
–
For this research, we analyzed only ChatGPT conversations from users who allow their data to be used to improve models. Before analysis, we removed account-linked identifiers and identifiable information, and we report only aggregate findings.
-
Traditional evaluations and deployment simulation for AI risk assessment
By
–
Traditional evaluations and red-teaming remain essential, especially for rare or severe risks. Deployment Simulation complements them by helping us estimate how often undesired behaviors may occur in realistic use and surface new behaviors before release.
-
Anthropic research tracks Claude Code usage, tasks, and expertise
By
–
Our latest economic research introduces a framework for tracking Claude Code as it scales. Who is using Claude Code, and what are they using it for? How is the value of tasks changing? And how much does domain expertise shape whether a session succeeds?
-

The jump from version 5.1 to 5.2 seems short in benchmarks
By
–
The jump from the previous version in several benchmarks like DeepSWE or HLE, the truth is it makes me think that the jump from 5.1 to 5.2 seems short to them.
-
Frontier AI recreates world’s literature as playable video games
By
–
using frontier ai to recreate the world's literature as playable video games ftw!
-

Two levels of reasoning high and max for long horizon
By
–
The model has two levels of reasoning, high and max, which push its reasoning and programming capabilities over a long horizon.
-

IneffableLabs builds super-learner with NVIDIA Vera Rubin NVL72
By
–
@IneffableLabs builds a super-learner that learns from experience, requiring enormous compute scale. They chose NVIDIA Vera Rubin NVL72, one of the largest clusters on @GoogleCloud, to power their mission to achieve superintelligence.
-
Evals overestimate ability in creative or out-of-distribution tasks
By
–
I think the evals overestimate ability in creative or out-of-distribution tasks.
-
Nemotron 3 Ultra and the Open Model Landscape broadcast
By
–
Nemotron 3 Ultra and the Open Model Landscape | Nemotron Labs https://
x.com/i/broadcasts/1
rGmqqeRorDGy
…