Executable Code Actions Elicit Better LLM Agents Wang et al.: https://
arxiv.org/abs/2402.01030 #ArtificialIntelligence #DeepLearning #MachineLearning
LLMS
-

Executable Code Actions Improve LLM Agent Performance
By
–
-

New Agentic Leaderboard Ranks LLMs for Agents
By
–
Our new Agentic leaderboard is now live! I've long wanted a way to quickly know which LLM is best for powering agents. So we've just built a leaderboard with Albert Villanova! This ranks LLMs powering a smolagents CodeAgent on subsets of various benchmarks. GPT-4.5
-
Copilot Launches Voice, Reasoning, and Mac App Updates
By
–
Busy last few weeks + more coming…
-free unlimited voice + reasoning
-Think Deeper runs on o3-mini high now
-voice speaks 40+ languages
-Mac app in a dock near you
-gaming coming to Copilot Labs
Tip of the iceberg. Can't wait to show you what else the team is working on -

Detecting Misbehavior in Frontier Reasoning Models
By
–
Detecting misbehavior in frontier reasoning models Chain-of-thought (CoT) reasoning models “think” in natural language understandable by humans. Monitoring their “thinking” has allowed us to detect misbehavior such as subverting tests in coding tasks, deceiving users, or giving
-
Anthropic Claude Surges to Match OpenAI Model Usage
By
–
Anthropic usage surged in 2024. Since Claude 3.5 Sonnet’s June 2024 launch, OpenAI and Anthropic models have shared nearly equal usage, highlighting growing competition among text-based LLMs. (2/6)
-
DeepSeek R1 V3 capture 7% market share open-weight models
By
–
DeepSeek-R1 and V3 went from no usage in December 2024 to capturing 7% of messages at their peak, a significantly higher level than any previous open-weight model family. (3/6)
-

AI Model Usage Trends: Early 2025 Ecosystem Analysis Report
By
–
New Report: How has AI usage changed over the last year? Our analysis of model usage on Poe provides insights around user preferences, adoption rate, and disruptions. Key highlights from our Early 2025 AI Ecosystem Trends report:
-
Mega prompt template resource for ChatGPT
By
–
Thanks for your interest. Here's the mega prompt link: https://
godofprompt.ai/chatgpt-mega-p
rompt-template?utm_source=twitter&utm_medium=giveaway&utm_campaign=lead-prompt-template
… -
AI21 Labs Prioritizes Enterprise Value in Model Development
By
–
We build and measure our models first and foremost according to the value they bring to the enterprise. Right now we see less of that in the enterprise. As these trends evolve to include additional players, we will of course update future benchmarks accordingly.