can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and how many might have failed) iautomatically generalize to everything else when there is not a shred of evidence that they do.
@garymarcus
-

LLM Hallucinations and Visual Misinterpretation Research
By
–
Yowza! Similar in some ways to the Stanford paper recently on LLMs hallucinating responses to images they never saw.
-

Addressing the replication crisis in biomedical machine learning research
By
–
Biomedical AI may be headed for a replication crisis. (This work below is not about AI-generated reports; it’s about studies of biomedicine that use ML in their methods, and how they are evaluted.)
-
Inquiry into AI model training methodology and data augmentation
By
–
i would be pretty interested to know if anything other than scale had changed and want to know about the training, whether there was a lot of symbolically generated augmented data, etc
-
Inquiry into the architectural basis of recent AI math breakthroughs
By
–
is the new math result neurosymbolic with Lean, harnesses etc or a pure LLM?
-
Neurosymbolic systems vs LLMs for mathematical reasoning
By
–
i am betting it was a neurosymbolic system rather than a pure LLM, and i have already said that’s a route to doing well in math. have not seen the details
-
Debating the Role of Neurosymbolic Systems in Mathematical AI
By
–
how much you want to bet that symbolic tools such as lean were involved and that this was not a pure LLM? i have said numerous times that neurosymbolic systems do well on math: pretty sure this was one.
-

Critique of AI Development and Integrity
By
–
two trillion dollars to build “pathologically dishonest” AI
-

New METR Study Highlights Critical Safety Failures in AI Agents
By
–
Breaking If we can’t make AI agents follow rules, we are screwed. New study from METR reports that “when the agents were faced with hard tasks, they routinely violated constraints” This—routine breaking of rules— is why in a nutshell we absolutely need a different
-
Discussion on Waymo’s transition to a foundation model
By
–
if waymo shows data that the Waymo Foundation Model is doing better on accuracy and generalizability, i will certainly take note. (or if Waymo states clearly an unambigiously that the Waymo Foundational Model has entirely displaced the prior system, which that blog does not say).