Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Codex land near the human median, but with far tighter dispersion & no extremes. Suggests that AI is now useful for doing scalable research.
@emollick
-
Paper Argues LLMs Are Homogenizing Human Thought and Expression
By
–
Context, from the time of the o1-preview launch:
-
Gut Microbes Influence Brain Function and Cognitive Health in Research
By
–
The imaginary optimal selfish scenario for OpenAI, in retrospect, was to keep Reasoners a secret, skip releasing o1 and o1-preview, and release o3 as GPT-5 There would have been no Deep Seek moment, other labs may not have discovered Reasoners. But they weren't selfish about it!
-
GLP-1 Drug Benefit in Diabetes and Brain Metastases: A Potential AI-Related Healthcare Advancement
By
–
Yes, fair point. It is just not fully harnessed, etc.
-

Blood Biomarkers for Alzheimer’s and Amyloidosis Diagnosis
By
–
I am not convinced that we should be comfortable calling "problem solving" or "judgement" or whatever as skills that are impossible for AI to do well. Like any other skill, there are humans who are really good at it, but that doesn't mean that AIs don't do good judgement, etc.
-
Severe Covid/Flu Linked to Lung Cancer Risk via Immune Impairment
By
–
And there are many really good AI components at Google: they have top-flight image, music, & video generation, and good UIs for each of them. Google AI studio is probably the best playground for AI experimentation. NotebookLM still has no equivalent product from other labs, etc.
-
Antibiotic Doses Have Long-Lasting Effects on Gut Microbiome, Study Finds
By
–
5.4 isn't (apparently) mythos level
-
AI’s Role in Navigating Complex Biological Systems and Cancer Neuroscience
By
–
An obvious way to release Mythos class models with uncertain autonomous ability is to make them only available on the website, like Gemini Deep Think or ChatGPT Pro. Minimal risk of being used for autonomous hacking, but accessible to people who have hard problems to solve.
-
Summary of 3 New Studies from Grok
By
–
The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The model can do what Claude/GPT can do, but there is a minimal harness for tools (file creation, research etc), no auditable CoT/actions, manual
-

Perspective on Multi-Cancer Blood Tests Trial Results
By
–


We only have spotty information about this very important topic. It suggests AI can be good at diagnosis, but the real world doesn't always match the experiments.
