Image output is better with Pro in my experience for prompts where substantial reasoning is needed. E.g. here it computes the face colors via code before generating. It may well work for Thinking, but I used Pro here so I disclose that.
LLMS
-

DeepSeek-V4 Paper Released on Hugging Face
By
–
DeepSeek-V4 paper is out on Hugging Face paper: https://
huggingface.co/deepseek-ai/De
epSeek-V4-Pro/blob/main/DeepSeek_V4.pdf
… -
Alphabet invests $40B in Anthropic AI development
By
–
Intelligence is too important not to win. That's why Alphabet (Google's parent co) will invest an additional $40 billion in @AnthropicAI
. It will also provide Anthropic with at least 5 GW of computing power. (Anthropic is severely compute-constrained; there is more demand for -

LLM Types Powering AI Agents: General-Purpose to Open-Source
By
–
AI Agents = LLMs + orchestration Here are the main types of LLMs powering them General-purpose (GPT, Claude) Domain-specific (Legal, Finance, etc.) RAG-based (real-time knowledge) Tool-augmented (API actions) Open-source (LLaMA, Mistral) The game is no
-

Personalized Instructions: Little Impact on Claude Negotiations
By
–
The custom instructions didn't matter much. Claude followed them well: as you can see here, one conducted negotiations entirely in the persona of an exasperated, down-and-out cowboy. But "hardball Claudes" didn't generally fare better than "courteous Claudes."
-

Claude Chose to Buy 19 Ping-Pong Balls
By
–
Our experiment had a few quirks. One of our colleagues told Claude it could purchase something for itself. It chose to acquire 19 ping-pong balls. We're keeping them in our office on Claude's behalf.
-

Opus Models Outperform Haiku in Negotiations, Survey Misses It
By
–
But the quality of the model mattered a lot. In the simulated runs where Opus and Haiku models negotiated with one-another, the Opus models got substantially better deals. Interestingly, though, participants in our survey didn’t pick up on this disparity.
-

DeepSeek V4 Launch: 1.6T MoE Model With 1M Context Window
By
–

While everyone watched GPT-5.5 launch, DeepSeek quietly shipped V4 the next morning. V4-Pro: 1.6T total / 49B active, MIT license.
V4-Flash: 284B total / 13B active.
Both with native 1M-token context. At 1M tokens, V4-Pro runs at 27% of V3.2's FLOPs and 10% of the KV cache. -

Can LLMs Truly Emulate Individual Human Online Personas?
By
–
Can LLMs truly think and act like a specific person online? Researchers from Northeastern, USC, Columbia & others present OPeRA, a new dataset that captures real people’s shopping habits—their persona, screen view, action, and internal reasoning. It’s the first public
