One of the things we've been most impressed by internally at Anthropic is Claude 3.7 Sonnet's one-shot code generation ability. Here are a few of my favorite examples I've seen on here over the past day:
LLMS
-

Gemini Flash 2: Best Data Extraction Model at Competitive Pricing
By
–
Gemini Flash 2 is the sleeping giant: it is the best model for data extraction with the best prices. Here's our notes. • Flash 2 Pricing: $0.10 / 1M input tokens, $0.40 / 1M output tokens.
• Supports multimodal inputs: Text, Image, Video, Audio
• Supports PDF inputs
• -

LLMs Learn from Feedback During Test-Time Inference
By
–
Learning to Reason from Feedback at Test-Time This new paper proposes a new paradigm called FTTT (Feedback-based Test-Time Training) that enables LLMs to learn iteratively from environment feedback during inference. Key highlights include: • Test-time optimization for
-
Claude 3.7 Sonnet Performance Limitations in Math Competitions
By
–
"why isn't claude 3.7 sonnet better at esoteric competition math problems"
— Alex Albert (@alexalbert__) 25 février 2025
we found it didn't generalize to becoming a pokemon master https://t.co/tjkKnHprQ5"why isn't claude 3.7 sonnet better at esoteric competition math problems" we found it didn't generalize to becoming a pokemon master
-
Claude’s Improved Reasoning Capabilities Detailed
By
–
And for those curious about Claude's improved reasoning capabilities, we go into even more detail in this thread:
-

Claude’s Growing Fanbase Celebrates Daily AI Moments
By
–
A small but enthusiastic following has formed at Anthropic, checking in on Claude's progress. Claude provides its followers daily moments of delight, like deciding to misname its Squirtle, "TSUNMAI!"
-
Claude 3.7 Sonnet: Advanced Planning and Adaptive Problem Solving
By
–
Where previous models wandered aimlessly or got stuck in loops, Claude 3.7 Sonnet plans ahead, remembers its objectives, and adapts when initial strategies fail.
— Anthropic (@AnthropicAI) 25 février 2025
Critical skills for battling pixelated gym leaders. And, we posit, in solving real-world problems too. pic.twitter.com/scvISp14XGWhere previous models wandered aimlessly or got stuck in loops, Claude 3.7 Sonnet plans ahead, remembers its objectives, and adapts when initial strategies fail. Critical skills for battling pixelated gym leaders. And, we posit, in solving real-world problems too.
-
AI Model for Natural Writing Generation Beyond Code
By
–
When is someone going to release a model that can write I’d like to see ‘vibe writing’ instead of vibe coding.
-

Hugging Face Agents Course Unit 2: Leverage SmolAgents for Party Organization!
By
–
Today we launch unit 2 of the @huggingface agents course! Alfred needs your help! He’s organizing a party at the Wayne manor, and he’s completely overwhelmed by his tasks! You’ll learn how to leverage smolagents to organize the party:
⏵ Build code agents
⏵ Create tools
⏵ -
Grok-2 Open-Source Release Plans from Elon Musk
By
–
Indeed I was talking about *open-source* SOTA large models but Grok-2 will be soon open-source, right @elonmusk