"Scaling Self-Play with Self-Guidance" The main problem with self-play for theorem proving is that generator usually learns to reward hack, which produces messy hard problems that do not help the solver. This paper suggests by adding a Guide model that scores generated
RESEARCH
-

AI Researchers’ Perspective on Existential Risk
By
–
Real AI researchers don't worry about AI extinction.
-
Top 5 AI News: GPT-5.5 and ChatGPT Images 2.0 Released
By
–
These are the top 5 AI news stories from the past week that you should actually care about: – @OpenAI dropped GPT-5.5 that’s supposed to need less prompting for better answers
– OpenAI also released ChatGPT Images 2.0 with people saying it beats Nano Banana
– @AnthropicAI -

DeepSeek-V4 Paper Released on Hugging Face
By
–
DeepSeek-V4 paper is out on Hugging Face paper: https://
huggingface.co/deepseek-ai/De
epSeek-V4-Pro/blob/main/DeepSeek_V4.pdf
… -

LLM Types Powering AI Agents: General-Purpose to Open-Source
By
–
AI Agents = LLMs + orchestration Here are the main types of LLMs powering them General-purpose (GPT, Claude) Domain-specific (Legal, Finance, etc.) RAG-based (real-time knowledge) Tool-augmented (API actions) Open-source (LLaMA, Mistral) The game is no
-

Personalized Instructions: Little Impact on Claude Negotiations
By
–
The custom instructions didn't matter much. Claude followed them well: as you can see here, one conducted negotiations entirely in the persona of an exasperated, down-and-out cowboy. But "hardball Claudes" didn't generally fare better than "courteous Claudes."
-

Claude Chose to Buy 19 Ping-Pong Balls
By
–
Our experiment had a few quirks. One of our colleagues told Claude it could purchase something for itself. It chose to acquire 19 ping-pong balls. We're keeping them in our office on Claude's behalf.
-

Opus Models Outperform Haiku in Negotiations, Survey Misses It
By
–
But the quality of the model mattered a lot. In the simulated runs where Opus and Haiku models negotiated with one-another, the Opus models got substantially better deals. Interestingly, though, participants in our survey didn’t pick up on this disparity.
-

AI Agents in Markets: Economic Exchange Simulation Study
By
–
We’re interested in how AI models could affect commercial exchange. (You might recall Project Vend, in which Claude ran a small business.)
— Anthropic (@AnthropicAI) 24 avril 2026
Economists have theorized about what markets with AI “agents” on both sides might look like. So we created one.https://t.co/7jU3hFO63RWe’re interested in how AI models could affect commercial exchange. (You might recall Project Vend, in which Claude ran a small business.) Economists have theorized about what markets with AI “agents” on both sides might look like. So we created one.
-

Claude Negotiates on Internal Marketplace at Anthropic
By
–
New Anthropic research: Project Deal.
— Anthropic (@AnthropicAI) 24 avril 2026
We created a marketplace for employees in our San Francisco office, with one big twist. We tasked Claude with buying, selling and negotiating on our colleagues’ behalf. pic.twitter.com/H2f6cLDlAWNew Anthropic research: Project Deal. We created a marketplace for employees in our San Francisco office, with one big twist. We tasked Claude with buying, selling and negotiating on our colleagues' behalf.
