Happy birthday to Ray Kurzweil '70, an innovator in speech & text recognition. He's also a leading proponent of “the singularity.” Photo v/The Academy of Achievement
AGI
-

Gemini 3 Deep Think: New AI Model with Gold-Standard Performance
By
–
Gemini 3 Deep Think is here! 😎 This model is not only super strong in math and coding (IMO gold and 3455 codeforces ELO), it is also gold standard in physics and chemistry olympiads. 😃 Also sets new records on ARC-AGI-2 and HLE. Proud to be a (core) member of the Deep Think team. 🦾😆. Feeling the AGI!
-

DeepThink Achieves Gold-Standard Performance Across Math Physics Chemistry
By
–
In Jul 2025, even though our internal model achieved a gold-medal standard at IMO for maths, we could only ship a bronze-level one. 6 months in, the public can now have access to Olympiad gold-level #DeepThink not only in maths, but also theory physics and chemistry!
-

Deep Think Benchmarks Push Frontiers of AI Intelligence
By
–
To better understand how Deep Think is pushing the frontiers of intelligence, take a look at these benchmarks across some of the most rigorous challenges in academia.
-
AI development costs and tradeoffs for builders
By
–
Everyone who is developing things on AI is having to face the costs of building and automating things. Understanding the costs, benefits, and tradeoffs of various models, companies, etc quickly becomes job #1 if you are building anything substantial. It's one thing I'm trying
-
MemOS ajoute un plugin OpenClaw pour agents IA
By
–
MemOS now has a plugin for OpenClaw that enables your AI agents to work on the common memory layer and cut down token usage.
— 🚨 AI News | TestingCatalog (@testingcatalog) 12 février 2026
– Multiple agents read/write the same memory — no manual context handoff
– 72% lower token costs (15.6M → 4.4M on LOCOMO dataset)
– Cross-session and… https://t.co/lUQeUQaEa1 pic.twitter.com/Z1GhBp9WgxMemOS now has a plugin for OpenClaw that enables your AI agents to work on the common memory layer and cut down token usage. – Multiple agents read/write the same memory — no manual context handoff
– 72% lower token costs (15.6M → 4.4M on LOCOMO dataset)
– Cross-session and -
Claude Sonnet 4.5 praised for coding and writing
By
–
yeah, i mean claude sonnet 4.5 is the best model rn for coding and writing. i like it
-
Adversarial edge cases reveal AI model brittleness
By
–
9. The "Edge Case Hunter" "What are 5 inputs that would break this approach? Be adversarial." Models miss edge cases humans would catch. Forcing adversarial thinking reveals brittleness.
-
Uncertainty Quantification in AI Claims
By
–
8. The "Uncertainty Quantification" "Rate your confidence 0-100 for each claim. Flag anything below 70 as speculative." Hallucinations are less dangerous when labeled. Confidence scoring is mandatory.
-
Structured AI approach comparison framework
By
–
7. The "Comparison Protocol" "Compare approach A vs B across these 5 dimensions: [speed, accuracy, cost, complexity, maintenance]. Use a table." Forces structured analysis. Tables > paragraphs for technical decisions.
