AI news not to miss from June 12, 2026: – New York imposes a label for AI actors in ads
– New York wants to freeze chatbot toys for five years
– Microsoft Scout transforms Copilot into an agent that acts without waiting
– Hyundai and NVIDIA want to test AI robots to
SAFETY
-

New York imposes AI label and freezes toys, Microsoft Scout, Hyundai-NVIDIA robots
By
–
-

Article against Anthropic as safety-branded permission regime for Cognition Infrastructure
By
–
Wrote an article on the case against Anthropic as a safety-branded permission regime for Cognition Infrastructure Covers sabotage-as-safety, anti-Opensource rules, Fable, regulatory capture, data asymmetry, Claude Code, and Who Owns Intelligence Must read, bookmark it for later
-
Anthropic needs certification department for powerful model users
By
–
I think Anthropic needs to build a certification department that audits and approves users of powerful models. Computer security companies, biotech researchers, academic labs, doctors, government institutions need access to the best AI we can build.
-
Fable project triggers safeguards, falls back to Codex
By
–
had an idea for a big fable project, set it up, and let it cook came back an hour later and it had triggered the safeguards and fell back to 4.8 10 minutes in back to codex
-

Yudkowsky praises paper on gradient descent difficulty and verification
By
–
On a first read, this paper seems far ahead of the pack in terms of (1) understanding some reasons why a task might stay difficult even in the face of gradient descent, and (2) distilling out propositions they'd need to somehow verify before they started expecting nice things.
-
First fast read: key point – absent method, ASI must not proceed
By
–
Good paper on a first fast read. I have misc quibbles, eg "setting aside" the chance that N can't align N+1; the more a fair solution is hard, the more likely a fake solution gets found instead. The main point not spelled out is, "Absent a method, ASI must not proceed."
-
Defending Anthropic using ToS vs training on all knowledge
By
–
To everyone defending Anthropic by mentioning how things are against their ToS: Cry me a river they trained their models on the entirety of humanity's knowledge
-

AI justification warning: post hoc reasoning not reliable
By
–
Here is the justification (but treat post hoc justifications with suspicion, since AIs are not able to reflect on their own thinking)
-

Frontier AI models fail to change ‘three words’ to ‘four’ when appropriate
By
–
This is an interesting test, and the frontier models (GPT-5.5 Pro Extended, Claude 5 Fable Max) do fail. They refuse to turn the "three words" into "four" if that fits better Prompting the AI to act like a translator surfaces the problem, but it still avoids changing the wording
-

Auto-review becomes default for new users with 97% accuracy
By
–
Auto-review is now the default for all new users. A classifier subagent reviews actions in context before deciding whether to allow, block, or ask for approval. Our evals show it's 97% accurate, with most misses near ambiguous edges.