Andrej Karpathy has launched a new blog to share his "random" notes. Add it to your RSS list—definitely worth following.
Link:
LLMS
-
Andrej Karpathy Launches New Blog Sharing AI Research Notes
By
–
-
GroupM Integrates DeepSeek R1 for Chinese Audience Targeting
By
–
Earlier today, @GroupMWorldwide said it's integrated @deepseek_ai
's R1 model into GroupM's Audience Translator to help advertisers find audiences in China. R1 improves NLP in Chinese, helps uncover consumer preferences and behaviors & turns audiences into addressable audiences. -

Brain Neural Activity Aligns With LLM Contextual Embeddings
By
–
Inspired by the success of LLMs, today on the blog we discuss how neural activity in the human brain aligns linearly with the internal contextual embeddings of speech and language within LLMs as they process everyday conversations. Learn more →
https://
goo.gle/4iiUoNj -
CollaborativeAgentBench: First Multi-Turn Human-Agent Collaboration Benchmark
By
–
New agents benchmark: CollaborativeAgentBench is the first benchmark studying collaborative LLM agents that work with humans across multi-turn collaboration on realistic tasks in backend programming & frontend design
-
What is Natural Language Processing?
By
–
What is Natural Language Processing or NLP? pic.twitter.com/U77TIU6C2M
— God of Prompt (@godofprompt) 21 mars 2025What is Natural Language Processing or NLP?
-
Discussion on potential improvements for AI model output
By
–
Yeah, I hope it is just the 1st version and not an actual grok 3 native output. If native will be released it it should be a massive improvement
-
Engineering at Anthropic: Developer Hub for Claude Optimization
By
–
This is our latest post in our new blog: Engineering at Anthropic. We've created this hub for developers to find practical advice and our latest discoveries on how to get the most from Claude.
-

Claude’s ‘Think’ Tool Enhances Agent Instruction Adherence
By
–
New research from our team at @AnthropicAI shows how giving Claude a simple 'think' tool dramatically improves instruction adherence and multi-step problem solving for agents. We've documented our findings in a blog post:
-
Influencing LLM Development Through High-Quality Evaluations
By
–
Most people don't realize they can significantly influence what frontier LLMs improve at, it just requires some work. Publish a high-quality eval on a task where models currently struggle, and I guarantee future models will show substantial improvement on it.
-
Claude’s Think Tool Advances Agentic Capabilities Remarkably
By
–
The latest post is about a new method, the “think” tool, that can result in remarkable improvements in Claude’s agentic tool use ability: