That’s true but probably the smallest asterisk on this because other models tested after only show relatively small gains. My guess is browsing and code execution each matter more
AGENTS
-
Grounding AI Hallucinations: Human Verification in Deep Research
By
–
with Deep Research, how are you going to ground hallucinations if not people helping ground them?
Maybe wikipedia the tool goes away, unclear, but its hard to see a bunch of people-centric subjective concepts grounded directly by AI. -
Offline Reinforcement Learning: Beyond Mimicking to Improvement
By
–
Offline reinforcement learning, where an agent tries to improve a behavior policy by observing another agent without actually playing, is a harder problem than it appears. The challenge isn’t to mimic the provided play, but to learn something better than what you have seen. The
-
Learning to Build Agents from 11x Example
By
–
If you want to learn how to build agents… 11x is a great example to learn from
-
Access DeepSeek r1 671B privately via Apollo OpenRouter
By
–
Use the full 671B version of DeepSeek r1 privately from your phone. Apollo with OpenRouter will find a provider based on uptime and speed.
— Aaron Ng (@localghost) 3 février 2025
Best way to try it without giving DeepSeek your info. pic.twitter.com/UBhBjDhLTiUse the full 671B version of DeepSeek r1 privately from your phone. Apollo with OpenRouter will find a provider based on uptime and speed. Best way to try it without giving DeepSeek your info.
-
OpenRouter launches reasoning traces with expanded inline CoT design
By
–
Reasoning traces with OpenRouter out now, with an expanded inline CoT design coming soon.
-

Building Agents: Expert Panel with Andrew Ng and Industry Leaders
By
–
we're getting a great group of folks together to talk about building agents andrew ng, adam d'angelo, amjad & michelle from Replit, shreya shankar (awesome LLM evals work), folks from agent-first companies like Cognition, 11x and Factory come join us!
-

OpenAI Launches Advanced Research Agent and Reasoning Model
By
–
OpenAI strengthens its AI capabilities with the launch of a new research-focused agent and a more powerful reasoning model, marking significant advances in both autonomous research and model performance. Here's your daily AI news briefing for February 3rd, 2025: OpenAI
-

DeepResearch’s Human-Like Research Behavior Including Reddit Exploration
By
–
the most human part of DeepResearch is that somehow, in any large and serious research task, it ends up reading reddit threads just like I would do
-

Merge Club: Replit-Built Guide for Microgrants and Fellowships
By
–
Amazing Replit Apps: Merge Club A web-based guide for discovering and understanding microgrant and fellowship opportunities worldwide. Merge Club is a seriously beautiful and well-designed site, built and deployed on Replit. From the creator: "I LOVE REPLIT AGENT!"