Our results suggest that CoT is less faithful on harder questions. This is concerning since LLMs will be used for increasingly hard tasks. CoTs on GPQA (harder) are less faithful than on MMLU (easier), with a relative decrease of 44% for Claude 3.7 Sonnet and 32% for R1.
LLMS
-

Chain-of-Thought Reasoning Lacks Faithfulness in AI Models
By
–
We found Chains-of-Thought largely aren’t “faithful”: the rate of mentioning the hint (when they used it) was on average 25% for Claude 3.7 Sonnet and 39% for DeepSeek R1.
-

Reasoning Models Don’t Reveal When Using Hints
By
–
We slipped problem-solving hints to Claude 3.7 Sonnet and DeepSeek R1, then tested whether their Chains-of-Thought would mention using the hint (if the models actually used it). Read the blog: https://
anthropic.com/research/reaso
ning-models-dont-say-think
… -

Anthropic Research: Reasoning Models Fail Verbalize Accurately
By
–
New Anthropic research: Do reasoning models accurately verbalize their reasoning? Our new paper shows they don't. This casts doubt on whether monitoring chains-of-thought (CoT) will be enough to reliably catch safety issues.
-
Convergence AI upgrades Deep Work with parallelisation
By
–
BREAKING 🚨: @convergence_ai_ released an upgrade to their Deep Work feature, which now supports parallelisation of agentic work. pic.twitter.com/RyKw98pVwg
— 🚨 AI News | TestingCatalog (@testingcatalog) 3 avril 2025BREAKING : @convergence_ai_ released an upgrade to their Deep Work feature, which now supports parallelisation of agentic work.
-

AMIE: Advancing AI for Longitudinal Disease Management
By
–
From diagnosis to treatment: Advancing AMIE for longitudinal disease management https://
buff.ly/eCdyiJZ
#AI #MachineLearning #DeepLearning #LLMs #DataScience -
Gemini 2.5 Pro Now Available on Poe Platform
By
–
Gemini 2.5 Pro is available at https://
poe.com/Gemini-2.5-Pro
-Exp
… and on Poe apps across all platforms. (2/2) -

Gemini 2.5 Pro Launches on Poe as World’s Most Powerful Model
By
–
Now on Poe: Gemini 2.5 Pro! This is the most powerful model in the world today according to a few benchmarks, and it is #1 ranked on LMArena by a large margin. It supports text, image, video, and audio input, and has a 1M token context window. (1/2)
-

Mixture of Routers: Novel Architecture for AI Model Routing
By
–
Mixture of Routers
Paper: https://
arxiv.org/pdf/2503.23362
Code: https://
anonymous.4open.science/r/MoR-DFC6 -

Mixture of Routers: Advanced MoE Architecture for DeepSeek
By
–
MoE is powerful and a key foundation for models like DeepSeek. SUES takes it further with Mixture of Routers (MoR), applying MoE to routers! MoR uses multiple subrouters for joint selection, with a learnable main router to weight them—and it performs impressively well.
