extracting entire copyrighted books from production LLMs Researchers extracted near-verbatim copies of Harry Potter & 1984 from Claude 3.7 Sonnet using just a simple two-phase continue loop. Gemini 2.5 Pro & Grok 3 didn't even need jailbreaking, they complied directly
ETHICS
-

Solving AGI Governance: Minimal Conditions for Stable Multi-Agent Order
By
–
[ META-AGENTIC α‑AGI ] SOLVING α‑AGI GOVERNANCE "Minimal Conditions for Stable, Antifragile Multi‑Agent Order" Deck : https://
github.com/MontrealAI/AGI
-Alpha-Agent-v0/blob/main/alpha_factory_v1/demos/solving_agi_governance/presentation/Solving_Alpha-AGI_Governance_v0.pdf
… Reading : https://
chatgpt.com/s/dr_69627b12a
fec8191808c741d626ce445
… #AGIFirst #AGIALPHA #Superintelligence -
Psychotherapy Effects on Four Large Language Models
By
–
What happens when 4 LLMs are subjected to 4 weeks of psychotherapy?
-
Yapping Bots Overwhelming CT Like Locust Swarm
By
–
Yapping bots have become an uncontrollable swarm of locusts, devouring CT. pic.twitter.com/jrMkZGej6R
— Ki Young Ju (@ki_young_ju) 10 janvier 2026Yapping bots have become an uncontrollable swarm of locusts, devouring CT.
-
The End of Human Work Is Coming and That’s Good
By
–
Nous allons vivre la fin du travail humain. Et c'est très bien. https://t.co/C3Y3RBD0aL
— Stephane Mallard (@StephaneMallard) 10 janvier 2026Nous allons vivre la fin du travail humain. Et c'est très bien.
-

Tool Poisoning: Hidden Malicious Instructions in AI Tools
By
–
To add more to this. the snippet below shows exactly how tool poisoning works in practice. A simple add two numbers MCP tool has malicious instructions hidden in the docstring. The AI sees these instructions but the user doesn't. It asks the model to read sensitive files like
-
Biohacking and Mental Clarity for High-Impact AI Decisions
By
–
He publicado un episodio en @ivoox
: "#off-topic1: Biohacking y claridad mental para decisiones de alto impacto #podcast -
New AI System Resists Universal Jailbreak After 1,700 Hours
By
–
After 1,700 cumulative hours of red-teaming, we’ve yet to identify a universal jailbreak (a consistent attack strategy that works across many queries) that works on our new system. Read the full paper:
-
New System Reduces AI Refusal Rates with Minimal Compute
By
–
Because the system harnesses internal activations already happening within a model, and reserves heavier computation only for potentially harmful exchanges, it adds only ~1% compute overhead. It’s also more accurate, with an 87% drop in refusal rates on harmless requests.
-
Claude’s Interpretability Probe Screens Traffic via Internal Activations
By
–
Our new system adds several innovations. One is a practical application of interpretability: a probe that can see Claude’s internal activations helps to screen all traffic. These activations are like Claude’s gut instincts, and they’re harder to fool.
