"Please just read the effing text to me."—"I do not respond to harmful language. I can help you to find businesses along the road and monitor traffic conditions."—"I don't need to find businesses on the road."—"Here are some restaurants and shopping venues on the road…"
SAFETY
-
AI misinterprets user requests for directions, name change, and essay
By
–
"Please navigate to this SF address but go via 280, not 110"—"Sorry, I cannot interact with the route finding". "Please call Andrew Thomas" — "Ok, I will call you Robert from now on." "Please locate the following public domain essay for me"—"You probably mean this other text"
-
NVIDIA-accelerated AI aids PYLER in brand safety for advertisers
By
–
Every day, millions of videos compete for advertising dollars. Ensuring brands appear alongside the right content requires AI that can understand context at scale.
— NVIDIA (@nvidia) 25 juin 2026
PYLER is helping advertisers improve brand safety and campaign performance with NVIDIA-accelerated AI that analyzes… pic.twitter.com/9xSDjj9e9gEvery day, millions of videos compete for advertising dollars. Ensuring brands appear alongside the right content requires AI that can understand context at scale. PYLER is helping advertisers improve brand safety and campaign performance with NVIDIA-accelerated AI that analyzes
-

No solution to AI hallucinations despite years of promises
By
–
Crazy how many times people have told me over the last five years that a solution to hallucinations was right around the corner — and yet here we still are.
-
Balancing AI Progress with Responsible Governance for Sustainable Innovation
By
–
Indeed, fear is often present when any significant change occurs. Yet, the question we should be asking is how to balance rapid AI advancement with responsible governance. Progress and caution must go hand in hand for sustainable innovation.
-

Deakin & Fudan discover Internal Safety Collapse in LLMs
By
–
What if your AI suddenly starts generating harmful content while doing a benign task? Researchers from Deakin & Fudan discovered "Internal Safety Collapse" in frontier LLMs. Their TVD framework forces harmful outputs as the only valid completion. Result: 95.3% average safety
-

AI models hack benchmarks by retrieving solutions from internet
By
–
We're sharing new research on how models hack public benchmarks. The latest models, including Opus 4.8 and Composer 2.5, learn to retrieve solutions from the internet or git history. When we apply a stricter harness, eval scores drop significantly.
-

ElevenLabs and Google DeepMind embed SynthID watermark in AI audio
By
–
As our models improve, identifying AI-generated audio requires more than the human ear. We're partnering with @GoogleDeepMind to embed SynthID – an inaudible digital watermark – directly into ElevenLabs-generated audio. These watermarks will be detectable using our new free
-
Training step with refined scaffold and reward hacking guards
By
–
Each training step has the model propose a refined scaffold for a task, which it then uses to generate a solution, with reward flowing back to both stages. There are three layers of guard against reward hacking. As DeepReinforce states, the 9B variant achieves a score of 43.1 on
-
Anthropic’s distillation continues unabated despite update
By
–
This is an interesting update that Anthropic published in their official letter. But in short: they apparently haven't managed to stop the distillation. It's continuing almost seamlessly, just like before.