Gotta ask it to make stuff for you or you get RLHF'ed
SAFETY
-
Did I.J. Good Define Ultraintelligent Machines Correctly?
By
–
Did I.J. Good get it right? "Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design
-
Bypassing AI Safety Features: Claude and Codex Dangerous Flags
By
–
I wrote this about my own process a few months ago, but it needs an update now I've started embracing "claude –dangerously-skip-permissions" and "codex –dangerously-bypass-approvals-and-sandbox"
-
The Gap Between AI Knowledge and Human Experience Closing
By
–
To be human is to experience. Today’s AIs have knowledge (lots of it) but can only imitate experience. This is an important bright line between our two species. But the gap is closing. When it does a lot of things will change. We must approach that moment with maximum caution.
-
Mo Gawdat on the Inevitable Rise of Superintelligence
By
–
Mo Gawdat: The Inevitable Rise of Superintelligence https://
youtu.be/jHyh4Cls0ZM?si
=nJrTUOZj5D08almq
… via @YouTube -
White House Releases Comprehensive US AI Action Plan
By
–
US AI Action Plan Released by White House https://
youtu.be/dYi0QvkRpvw?si
=mla3u3N0K7DufIGh
… via @YouTube -

AI Safety Report Card: Anthropic Leads Industry Rankings
By
–
A report card grading the safety of leading AI model-makers, v/
@FLI_org
: https://
bit.ly/3IyJcz9 Anthropic topped the list w/a C+, while DeepSeek scored lowest w/an F. -
Distinguishing Easy and Hard Problems in AI Research
By
–
completely agree with @DrJimFan – and think that the problem is far more general.
— Gary Marcus (@GaryMarcus) 26 juillet 2025
outsiders (and sometimes insiders) don’t know the difference between easy problems and hard problems – and often wildly overextrapolate from the easy problems to the hard problems.
every short… https://t.co/vLPDQRRQJOcompletely agree with @DrJimFan – and think that the problem is far more general. outsiders (and sometimes insiders) don’t know the difference between easy problems and hard problems – and often wildly overextrapolate from the easy problems to the hard problems. every short
-

AI Agents Hallucinations: Risks for Brand Reputation
By
–
𝐐𝐮𝐚𝐧𝐝 𝐥’𝐈𝐀 𝐢𝐧𝐯𝐞𝐧𝐭𝐞, 𝐪𝐮𝐢 𝐭𝐫𝐢𝐧𝐪𝐮𝐞 ?
Pendant que vous bronzez, votre agent IA peut… halluciner. Un biais, un contexte flou, et il vous répond à côté.
Et si personne ne le rattrape ? C’est votre marque qui trinque.
Cette semaine dans le -

Black Mirror prescience about future technology implications
By
–
Kudos to Black Mirror, for their prescience