How good are our AI research agents at producing genuinely insightful reports? University of Science and Technology of China and Metastone Technology just released "Deep Research Bench II". This new benchmark features 132 complex research tasks across 22 fields, evaluating AI
RESEARCH
-
Exploit Costs More Than Human Security Researchers
By
–
-
Anthropic and OpenAI: Decoding Strategic AI Model Names
By
–
At face value, the claim sounds like a safety warning. It also seems to signal something unusually capable, consequential and maybe beyond what rivals have. It’s hard not to notice the names @AnthropicAI and @OpenAI are using:
— "Spud" (still just a codename) calls to mind -
Advice to VCs: Avoid Technology Dependency in Robotics
By
–
Advice to VCs investing in robotics: when the underlying technology landscape is still changing SO incredibly quickly, avoid making bets that make you path dependent on a specific technology approach. (e.g. Vision Language Action Models vs. @GeneralistAI's System1/System 2 approach). We're still so early and have only seen ~0.1% of the ideas and innovation in this space. Invest in infrastructure & picks and shovels. [Translated from EN to English]
→ View original post on X — @ken_goldberg, 2026-04-13 18:39 UTC
-

Anthropic Mythos Sets New AI Competition Bar for OpenAI
By
–
Now that @AnthropicAI has introduced Mythos, the bar is set for @OpenAI and “Spud.” In today's @BigTechnology newsletter, @Kantrowitz look at what comes next. We also ask series of questions — like whether “too dangerous” is a new trend that's part warning & part marketing.
-
FLOPS: From Rate Measurement to Algorithm Work Quantification
By
–
FLOPS was originally “floating point operations per second”, specifying a rate of work for a system: A SPARCstation 2 gave 4.2 MFLOPS. Today you also see it used as “floating point operations” for an algorithm, or an amount of work: This layer takes 8 GFLOPS.
→ View original post on X — @id_aa_carmack, 2026-04-13 18:35 UTC
-
UK AI Security Institute Evaluates Claude Mythos Preview Safety
By
–
Very interesting evaluation from the UK’s AI Security Institute of the not yet publicly available Claude Mythos Preview. On the happy side, in its current form, Myth is nowhere near as scary as Tom Fridman (who worries about schoolchildren accidentally taking down power grids)
-
Gary Marcus evaluates Claude AI myths and capabilities
By
–
nice work! shared it here, please DM: https://
open.substack.com/pub/garymarcus
/p/claude-mythos-evaluated?r=8tdk6&utm_medium=ios
… -
Modern AI Models Distinguish Useful Links From Product Promotion
By
–
nah, modern models are smart enough to understand useful links vs unasked product shilling
-

Inverse Design Transforms Materials Discovery Process with AI
By
–
What if materials discovery started with the spec, not the search? It can with Domino. See how inverse design changes the process. Watch the demo: https://
hubs.ly/Q04bsJR-0
