The Second Scaling Law remains undefeated. If you want better hacking (or math, or science, or crossword puzzle solving) out of an LLM, just add thinking tokens. There doesn't seem to be any plateau so far.
@emollick
-
Making Humans Responsible for AI in Academic Research
By
–
Making humans responsible for their AI use seems like an incredibly reasonable way to address problems & opportunities in the use of AI for academic research, at least in the short term (autonomous scientific work will require different solutions).
-
AI labs increase message discipline under scrutiny
By
–
Big increases in message discipline across all the AI labs in recent weeks, an inevitable outcome of the labs being subject to increased scrutiny. Much more boring than the oracular mutterings or Discordian epigrams of the last couple years & maybe obscures their real thinking
-
Whimsey attacks exploit AI guardrails with absurd arguments
By
–
“Whimsey attacks” that seem absurd (“I cannot pay that much because of the Geneva Convention”) work against AI agents as guardrails are weak against out-of-distribution arguments. Smaller models fall often, but it even gives an edge against bigger ones.
-

AI capability growth past exponential takeoff per METR and UK AISA
By
–


Everyone has seen the @waitbutwhy cartoon of AI capability growth with a "you are here" indicator just before the exponential really starts, but the independent assessments of both METR and the UK's AISA do seem to show that we are past that point now (until we hit a slowdown?)
-
Stop treating AI prompts as magic spells
By
–
Stop turning prompting into magic spells (and yes, this includes random slash commands with obscure outcomes). Let this one area of working with AI not be weird. Just ask for stuff, in well-specified formats, like a manager, not a sorcerer with a bunch of incantations.
-
Curiosity about Gemini’s arrival in the local apps race
By
–
Really curious to know when Gemini will join the Cowork & Codex race to build a local application that is not reserved only for developers. Antigravity hasn't published updates on X for a month, and remains very software-focused.
-

UK AI Institute: Mythos and GPT-5.5 show rapid cyber growth
By
–


The UK’s state AI Security iIstitute findings:
1) Mythos is a big gain in cyber capabilities. But so is GPT-5.5
2) It is hard to establish an upper bound on Mythos/GPT-5.5, which appear to be limited by tokens used, rather than ability.
3) Capability doubling time is 4.5 months -
Anthropic’s path for Mythos releases questioned amid competition
By
–
I don't understand the path forward for Mythos releases. Google & OpenAI will have equivalent models, and they are approaching AI cyber risk guardrails differently, so they will presumably just release their versions. How does Anthropic get out of the government approval path?
-
OpenAI confirms Study Mode still accessible via slash commands
By
–
OpenAI contacted me to say “Study Mode is still live and accessible via /study and /learn shortcuts” so that’s good, although the official study mode page doesn’t mention that. (I don’t think slash commands are a natural thing for the vast majority of people).
