maybe. See the quote from the Eli Lilly guy in my substack yesterday, though: current AI is really not good at biological reasoning. “AI is far from curing cancer and most other diseases.] If you just ask them to solve biology or chemistry questions, they’re not particularly good
HEALTHCARE AI
-
AI’s Role in Advancing Brain-Computer Interfaces for Vision
By
–
Helping the blind to see. @ScienceCorp_ has very unusual brain/computer interface that someday we might all have, I hear. Will take a while, maybe a decade more of science. But this shows another good thing AI is doing.
-

Retrieval, Not Writing, Boosts AI Memory Accuracy
By
–
most people building AI agents obsess over how they WRITE memories.
turns out that's basically irrelevant. new research analyzed 9 different memory systems across 1,540 questions.
the finding? – retrieval method drives 20-point accuracy swings.
– write strategy? 3–8 points max. -

EmotionThinker: reasoning-based speech emotion recognition
By
–
What if your AI could explain exactly why it thinks you sound frustrated or happy? Researchers from The Chinese University of Hong Kong and Microsoft present EmotionThinker to solve this. They shifted speech emotion recognition from simple labels to deep reasoning. By using a
-

#CamFest: AI and the Future of Public Health Event
By
–
Our #CamFest event is just weeks away! Join Henrietta Hughes, Alastair Denniston, @lawrennd, Melanie Ivarsson & Rozelle Kane as they explore the challenges & opportunities of using #AI to create healthier communities for all. Grab your ticket now!🎟️➡️ eventbrite.co.uk/e/ai-and-th…
-
Models Know Longer Isn’t Always Better
By
–
the deeper point here connects to something the field keeps rediscovering. we trained reasoning models to think longer. then we discovered longer doesn't mean better. now this paper shows the models themselves already know that. they're generating stop signals that our inference
-
DeepSeek and Qwen3 performance improvements
By
–
specific numbers worth sitting with: > DeepSeek-R1-7B on MATH-500: 93% accuracy (up from 91.6%), tokens cut from 3,871 to 2,141 > DeepSeek-R1-1.5B on AIME 2025: accuracy jumps 6.2 percentage points > Qwen3-8B: response length halved from 18,342 to 9,183 tokens with no accuracy
-
SAGE: Efficient Reasoning with Confidence Checks
By
–
their solution: SAGE (Self-Aware Guided Efficient Reasoning). instead of generating token by token, SAGE extends chains in whole reasoning steps. after each step, it checks: is the model confidently signaling it wants to stop? if yes, reasoning ends. no fine-tuning. no new
-
Overthinking harms accuracy in AI responses
By
–
and it's not just wasted compute. overthinking actively hurts accuracy. DeepSeek-R1 produces responses 5x longer than Claude 3.7 Sonnet on AIME 2025 with comparable accuracy. QwQ-32B scores 2 percentage points HIGHER with its shortest answers using 31% fewer tokens. 72% of
-
RFCS Metric Reveals Early Correct Steps
By
–
first, the problem quantified. the researchers created a metric called RFCS (Ratio of First Correct Step) that tracks where in a chain of thought the correct answer first appears. on MATH-500, across every model tested, the right answer shows up well before the end in over half