IMO-Bench also includes #AnswerBench, a set of 400 problems with verifiable answers carefully chosen from past Olympiad competitions. The problems span across four IMO categories (Algebra, Combinatorics, Geometry, and Number Theory), were altered by experts to avoid memorization,
AGENTS
-

MLflow-Powered Domain-Specific Judges for Agent Bricks
By
–
Building accurate, trustworthy agents in Agent Bricks starts with better judges — because AI agents can only be as good as the ones evaluating them.
— Databricks (@databricks) 4 novembre 2025
We’re excited to streamline creation of domain-specific judges with new MLflow-powered capabilities in Agent Bricks:
•… pic.twitter.com/vEDE2Rabg6Building accurate, trustworthy agents in Agent Bricks starts with better judges — because AI agents can only be as good as the ones evaluating them. We’re excited to streamline creation of domain-specific judges with new MLflow-powered capabilities in Agent Bricks:
• -

Appreciation for Grok’s Direct Correction Approach
By
–
I appreciate how Grok doesn’t sugar coat corrections.
-

IMO Medalists Grade AI Homework with Human Verification
By
–
We do have teachers (IMO medalists) to grade the homeworks too 🙂 See the paper https://
arxiv.org/abs/2511.01846 where we recommend to augment with human verifications. -

MiniMax-M2 Open Source Model Available on Poe
By
–
MiniMax-M2 is now available on Poe! This open source model has a 200k token context window, has 230b parameters with an MoE architecture, and excels at coding and agent workflows. (1/2)
-
Comprehensive AI Agent Security Book by Ken Huang Chris Hughes
By
–
Ken Huang and Chris Hughes have done is all a great service here by capturing the most comprehensive book on AI Agent security available today. Will help us all accelerate artificial intelligence into service for humanity. pic.twitter.com/6DBCsghnKy
— Bob Gourley – e/acc (@bobgourley) 4 novembre 2025Ken Huang and Chris Hughes have done is all a great service here by capturing the most comprehensive book on AI Agent security available today. Will help us all accelerate artificial intelligence into service for humanity.
-

ProofAutoGrader: Automatic IMO Proof Evaluation Using Gemini
By
–
While human expert evaluation remains the gold standard for mathematical proofs, its cost and time intensity limit scalable research. To address this, we built #ProofAutoGrader, an automatic grader for IMO-ProofBench. The autograder leverages Gemini 2.5 Pro, providing it with a
-
LangChain Releases Human-in-the-Loop Agent Middleware Series
By
–
Over the next few weeks we'll be releasing a series of deep dive videos into our new prebuilt agent middlewares. We're starting off with one the most popular middlewares, human-in-the-loop! Require approval on sensitive tool calls before execution in just 1 LOC!
-
LangSmith for Improving Agent Quality at Scale
By
–
Why we built LangSmith for improving agent quality
— LangChain (@LangChain) 4 novembre 2025
As more agents move into production, teams need to move beyond vibe-checking and bring rigor to how they understand agent behavior at scale.
In this video, the LangSmith engineering team and Harrison (@hwchase17) sit down to… pic.twitter.com/q7jrXfWJo1Why we built LangSmith for improving agent quality As more agents move into production, teams need to move beyond vibe-checking and bring rigor to how they understand agent behavior at scale. In this video, the LangSmith engineering team and Harrison (
@hwchase17
) sit down to -

ALPHA Agent AGI Ethereum Launch Announcement
By
–
[ 1.ALPHA.AGENT.AGI.ETH ] Website : https://
agialpha.com $AGIALPHA #AGIALPHA
