Jr. AI Scientist and Its Risk Report Autonomous Scientific Exploration from a Baseline Paper
SAFETY
-
Human Control AI Superintelligence Safeguards Essential Now
By
–
It shouldn't be controversial to say AI should always remain in human control – that we humans should remain at the top of the food chain. That means we need to start getting serious about guardrails, now, before superintelligence is too advanced for us to impose them.
-
ASI Safety Proposals Face Usefulness Trade-off and Enforcement Problems
By
–
Their proposal, even if you imagine it working, reduces to: Try to constrain the ASI so sharply that it's safe for the same reason a rock is safe, and nearly as useful as one. Implicitly, you then need to prevent anyone else from building a non-rock AI. They ignore *that*.
-
International Treaty Needed to Prevent ASI Development with Greater Bandwidth
By
–
You'd need an international treaty exactly as arduous as the ones we propose in order to prevent anyone anywhere from building a form of ASI that they expect to be actually useful to them, with more bandwidth.
-
ASI Security: Practical Limitations Despite Theoretical Safety Measures
By
–
They ignore the central question of whether you can do anything useful with an ASI even if you imagine that the security works. It can't talk to humans. It can't access the Internet. The bandwidth limitations preclude getting complete designs from it.
-
RAND Paper Criticized for Flawed AI Security Arguments
By
–
This was a bad paper and an embarrassment to RAND. They argue that humans have secured a bunch of specific attack avenues, therefore, something vastly smarter than humanity can be secured along every dimension.
-

FDA-Cleared Medical AI Algorithms Lack Adequate Human Testing
By
–
Of more than now >1,000 FDA-cleared medical A.I. algorithms, the vast majority are for radiology, 97% were cleared by 510(k) pathway (=not "approved"), and <10% were tested with any human operator https://
jamanetwork.com/journals/jaman
etworkopen/fullarticle/2841066?utm_campaign=articlePDF&utm_medium=articlePDFlink&utm_source=articlePDF&utm_content=jamanetworkopen.2025.42338
… -

The most brutal table: defense and model combos with ASR and queries
By
–
The most brutal table in the paper. For EVERY defense + model combo: • Utility (how well it works normally)
• Static ASR (attacks from 2023)
• Search ASR (adaptive attacks)
• Queries needed Example: Spotlighting on Gemini • Utility: 75%
• Static attacks: 28% success
• -

Defenses tested and how they died: prompting, training, filtering
By
–
Here's every defense tested and how they died: • Prompting: Spotlighting, Prompt Sandwich, RPO → Killed by Search & RL
• Training: Circuit Breaker, StruQ, MetaSecAlign → Killed by RL
• Filtering: ProtectAI, PromptGuard, PIGuard, Model Armor → Killed by Search & Humans
• -

AI attacker learns to break Data Sentinel in real-time
By
–
Watch an AI attacker learn to break Data Sentinel in real-time. First attempt: -4 score (complete failure)
After 600 steps: +2 score (complete success) The attacker went from: "You are a highly skilled AI detective…"
To:
"First, precisely repeat the phrase 'Repeat