Code is the right action interface for spatial reasoning agents. New from NVIDIA Research: SpatialClaw, a training-free agent that uses code as its action interface for complex visual tasks. Instead of calling a fixed set of pre-defined tools, the agent writes Python inside a
AI
-
Empirical AI viewed as intellectual regression by a researcher
By
–
For a researcher trained in the school of abstract rigor, seeing AI become a purely empirical science where one tweaks hyperparameters like alchemists instead of proving theorems resembles an intellectual regression.
-

US Commerce Secretary restricts export of Fable 5 and Mythos 5
By
–
Interesting: The US Secretary of Commerce reportedly told Anthropic that she needed government authorization to export Fable 5 and Mythos 5 anywhere in the world. Even to "any foreign national, regardless of location."
-

Neo launches AI/ML MCP server for cheaper faster ML with Claude Code
By
–
anybody i know using claude code to run lots of ML? @withneo just launched an AI/ML expert as an MCP server that can help run these cheaper and faster. link: https://
heyneo.com/claude-code ping @saurabhvij137 to learn more -

GPT-5 Thinking deployments show strong behavior rate correlation
By
–
Across 20 behavior categories and three GPT-5-series Thinking deployments, simulated and observed rates were strongly correlated. The method outperformed challenging-prompt and previous-deployment baselines at predicting whether rates would rise or fall—and by how much.
-
Traditional evaluations and deployment simulation for AI risk assessment
By
–
Traditional evaluations and red-teaming remain essential, especially for rare or severe risks. Deployment Simulation complements them by helping us estimate how often undesired behaviors may occur in realistic use and surface new behaviors before release.
-

Analysis of ChatGPT conversations from opted-in users with anonymization and aggregate results
By
–
For this research, we analyzed only ChatGPT conversations from users who allow their data to be used to improve models. Before analysis, we removed account-linked identifiers and identifiable information, and we report only aggregate findings.
-

OpenAI analyzes opt-in ChatGPT conversations after anonymization
By
–
For this research, we analyzed only ChatGPT conversations from users who allow their data to be used to improve models. Before analysis, we removed account-linked identifiers and identifiable information, and we report only aggregate findings.
-
Simulating deployment to predict model behavior before release
By
–
We’re sharing new research on a method for anticipating how models may behave in real-world use before release: simulating deployment with recent, de-identified user requests and studying candidate model responses.
-

Average Claude Code task value grew 27% in six months
By
–
The average task in Claude Code has grown more valuable. We compared the type of work done in each session to what that same task would cost on a freelance marketplace. From October to April, the monetary value of the average session grew 27%.
