The irony. Using AI (GPT 5.5), errors have been identified that would affect 1/3 of the problems in the FrontierMath benchmark Tiers 1-4. This could represent a shift in the evaluations of AI's mathematical capabilities, which may be underestimated.
RESEARCH
-

Multi-Agent Synergy for Scaling Test-Time Compute
By
–
TMAS Scaling Test-Time Compute via Multi-Agent Synergy
-

Rebellious Student: Reversing Teacher Signals
By
–
Rebellious Student Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
-

Multi-Agent Synergy for Scaling Test-Time Compute
By
–
TMAS Scaling Test-Time Compute via Multi-Agent Synergy
-

AGI ALPHA: Scalable Substrate for Intelligence Organizations
By
–
"AGI ALPHA: A Scalable Substrate for Intelligence Organizations" Vincent Boucher, President of http://
MONTREAL.AI and http://
QUEBEC.AI Paper: https://
github.com/MontrealAI/agi
alpha-first-real-loop/blob/main/docs/manuscript/AGI_ALPHA_Unified_Publication_Final.pdf
… #AGIAlpha #QuebecAI #SovereignAI -

Scal3R: New AI Research for Neural 3D Scene Reconstruction from Video
By
–
What if AI could reconstruct kilometer-scale 3D scenes from video as effortlessly as humans perceive a room? Researchers from Zhejiang University and Horizon Robotics present Scal3R. They introduce a neural context memory that compresses long-range scene info, adapting on the
-

Soohak: A Benchmark for Evaluating Research-level Math in LLMs
By
–
Soohak A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
-
Sharing Papers and Apps on Hugging Face
By
–
Paper:
https://huggingface.co/papers/2605.10922
…
App:
https://huggingface.co/spaces/TencentARC/Pixal3D
…

