That paper is nearly a year old at this point, has there been any follow-on research that supports or refutes it?
AGI
-
Base Models Unreliable Self-Assessment Capabilities
By
–
That said, asking a base model directly about its own capabilities feels even less useful to me than asking an instruction-tuned model – models have always been inherently unreliable when it comes to answering questions about themselves
-

MMLU benchmark validity questioned: score improvements may not indicate linear capability gains
By
–
It is, by the way unclear if MMLU is a good measure of anything or if an improvement in scores that is linear is a linear increase in capabilities. That is, going from 89-90 may be a bigger deal than 79-80.
-
Are LLMs Really Seeds of Their Own Destruction?
By
–
Right, lots of people really want to believe that LLMs are the inevitable seeds of their own self-destruction – it's a very tempting narrative! I'm trying to understand if it's actually playing out that way
-
How large model developers approach AI risk mitigation
By
–
I see lots of people outside of those organizations talking about this as a problem Presumably the people training the largest models have been thinking pretty hard about how much of a risk this is and what mitigations they can put in place What are their thoughts?
-
Admission Essays No Longer Useful in AI Era
By
–
I deleted the original tweet because I don’t want this sort of willful misinterpretation of what I was writing about (which was admission essays are no longer useful) to spread.
-
AI Agents Rapidly Improving: Not Just Hype but Strategic Priority
By
–
I think seeing agents as "hype" is going to result in people being blindsided as they become more powerful in the coming months. Agents are the explicit goal of OpenAI & Google AI. Plus, the improvement in agents over the last year has been pretty rapid.
-
Future AI Threat: When Humans Truly Fear Machines
By
–
The Butlerian Jihad will start the moment people genuinely feel threatened by the machines. There's no risk of this happening with this generation of skin-deep human-mimicry bots, but eventually, one day, it will happen.
-
Devin AI Agent Report Questioned as Possible Prank
By
–
Given that it was working fine before and it doesn’t use WP, and that this appears to be a report, not automated, I am guessing prank/dislike of Devin
-

AI interaction feels magical but not quite perfect yet
By
–
Honestly it feels pretty magical to use, like interacting with a person. Not quite smart enough to pull it off yet all the time, but close.