o3-mini-high helped a Brookhaven National Laboratory researcher find novel exact solutions to a physical model: https://
arxiv.org/pdf/2503.23758
LLMS
-

AI Model Discovers Novel Physics Solutions at Brookhaven Laboratory
By
–
-
New slow HQ mode announced for low latency performance
By
–
we like low latency, give it a shot! or use the 'slow' hq mode (which is our main announcement)
-
CSE 291: AI Agents Course at UCSD
By
–
CSE 291 – AI Agents – UCSD CSE https://
buff.ly/QZcn4I5
#AI #MachineLearning #DeepLearning #LLMs #DataScience -

CURIE: Scientific Long-Context Understanding Benchmark for LLMs
By
–
We introduce CURIE, a scientific long-Context Understanding, Reasoning and Information Extraction benchmark to measure the potential of large language models in scientific problem-solving and assisting scientists in realistic workflows. Learn more at https://
goo.gle/4jah5Ds -

DeepSeek V3 Ranked 8th and 12th on SEAL Leaderboards
By
–
Narrative Violation—DeepSeek V3 is a competitive, but NOT top model. SEAL leaderboards have been updated with DeepSeek V3 (Mar 2025). – 8th on Humanity’s Last Exam (text-only).
– 12th on MultiChallenge (multi-turn). View the full rankings: http://
scale.com/leaderboard -

ArtificialAnalysis Benchmarks Help Developers Choose Best Models
By
–
@ArtificialAnlys benchmarks help devs pick the best models and providers for their workloads – s/o to the team there for helping the community build fast! If you use AA, consider being a part of their survey
-
Quick Lazy Prompting: When Imprecision Works for LLMs
By
–
Contrary to standard prompting advice that you should give LLMs the context they need to succeed, I find it’s sometimes faster to be lazy and dash off a quick, imprecise prompt and see what happens. The key to whether this is a good idea is whether you can quickly assess the
-
Model Context Protocol: Origins, Agents, and Trust
By
–
The Creators of Model Context Protocol with @dsp_ and @jspahrsummers
! https://
latent.space/p/mcp We asked ALL your burning questions:
– The Origin Story of MCP
– MCP vs OpenAPI
– Building Agents with MCP
– How many MCPs is too many?
– Authorization and Trust in MCP Servers
– -
Making Chain-of-Thought Monitoring Viable for AI Safety
By
–
To make CoT monitoring a viable way to catch safety issues, we’d need a way to make CoT more faithful, evidence for higher faithfulness in more realistic scenarios, and/or other measures to rule out misbehavior when the CoT is unfaithful. Read the paper: https://
assets.anthropic.com/m/71876fabef0f
0ed4/original/reasoning_models_paper.pdf
… -

Outcome-Based Training Improves Model Faithfulness Modestly
By
–
Does outcome-based training increase faithfulness? Only to a small extent. Training models to use their CoTs more effectively does make them more faithful, but the benefits quickly plateau.