We've addressed the quirks from previous models head-on. Significantly reduced reward hacking in code generation. Better instruction following. Less overeager responses. These models do what you ask, how you ask.
PROMPT ENGINEERING
-

Claude 4 Alternates Between Thinking and Tool Calls
By
–
Claude 4 is able to alternate between thinking and tool calls as well, in both claude dot ai and the API. We’ve seen this greatly improve Claude’s performance and tool calling ability.
-
Claude 4 Extended Focus: Full Day Workflow Demonstration
By
–
But it's not just coding.
— Anthropic (@AnthropicAI) 22 mai 2025
Claude 4 models operate with sustained focus and full context via deep integrations.
Watch our team work through a full day with Claude, conducting extended research, prototyping applications, and orchestrating complex project plans. pic.twitter.com/4kPpDiKf2aBut it's not just coding. Claude 4 models operate with sustained focus and full context via deep integrations. Watch our team work through a full day with Claude, conducting extended research, prototyping applications, and orchestrating complex project plans.
-

Claude Opus 4 and Sonnet 4: Hybrid Models with Extended Thinking
By
–
Claude Opus 4 and Sonnet 4 are hybrid models offering two modes: near-instant responses and extended thinking for deeper reasoning. Both models can also alternate between reasoning and tool use—like web search—to improve responses.
-

Mistral Launches Document AI: OCR-Powered End-to-End Solution
By
–
Meet Document AI, our end-to-end document processing solution powered by the world’s best OCR model! https://
mistral.ai/solutions/docu
ment-ai
… -

Veo 3 vs Wan 2.1: Physics Performance on Complex Video Prompts
By
–
Tried this prompt with Veo 3:
— fofr (@fofrAI) 21 mai 2025
> televised footage of a cat is doing an acrobatic dive into a swimming pool at the olympics, from a 10m high diving board, flips and spins, there is commentary (not slow motion)
Wan 2.1 still has the best physics for out-of-distribution prompts. https://t.co/SFuzWufSTW pic.twitter.com/S2F1HvtO0hTried this prompt with Veo 3: > televised footage of a cat is doing an acrobatic dive into a swimming pool at the olympics, from a 10m high diving board, flips and spins, there is commentary (not slow motion) Wan 2.1 still has the best physics for out-of-distribution prompts.
-

Imagen-4 Long Prompts Generate Superior Comic Illustration Outputs
By
–
I haven't played with Imagen-4 a lot yet, but it seems like long prompts give the best outputs. Here's a Reve prompt on Imagen-4: > A comic book illustration features a character with spiky hair and intense glowing eyes, wearing torn punk-inspired clothing adorned with chains
-

Veo 3 Requires Highest Quality Setting Default Flow
By
–
PSA: Flow defaults to Veo 2
You need to use "Highest Quality" for Veo 3 -
Reinforcement Fine-Tuning LLMs with GRPO Short Course
By
–
New Course: Reinforcement Fine-Tuning LLMs with GRPO!
— Andrew Ng (@AndrewYNg) 21 mai 2025
Learn to use reinforcement learning to improve your LLM performance in this short course, built in collaboration with @Predibase, and taught by @TravisAddair, its Co-Founder and CTO, and @grg_arnav, its Senior Engineer and… pic.twitter.com/j5AXn3swADNew Course: Reinforcement Fine-Tuning LLMs with GRPO! Learn to use reinforcement learning to improve your LLM performance in this short course, built in collaboration with @Predibase
, and taught by @TravisAddair
, its Co-Founder and CTO, and @grg_arnav
, its Senior Engineer and
