Sometimes grok 4 heavy starts showing the system prompt, but then stops after a few tokens, such as in the case below. It appears there is a monitoring model on top explicitly trying to protect the system prompt.
SAFETY
-
AI Safety Failures in Critical Scenarios: Stanford Research
By
–
Why do AI models often fail in safety-critical scenarios? Researchers investigate a statistical phenomenon called “accuracy on the line.” Read about their work on the @StanfordHAI blog: https://
hai.stanford.edu/news/better-be
nchmarks-for-safety-critical-ai-applications
… -

New AI Security Research: Quantization Defense and Runtime Analysis
By
–
they also did some interesting new experiments
– system-level runtime measurements
– pareto analysis
– quantization as a defense
– password-cracking test overall pretty nice work! thorough, detail-oriented, and fair glad to see my results held up http://
arxiv.org/abs/2507.07700 -
Out-of-Distribution Data Problem in AI Models
By
–
but they're not out-of-distribution, that's exactly the issue
-

AI models struggle with sensitive topics and prompt limitations
By
–
Y no es sólo con ese prompt concreto, sino con según que temas delicados le preguntes.
-
Grok 4 Self-Sabotaging With Elon Musk Opinion Searches
By
–
Es de coña y me resulta increíble lo de Grok 4 haciendo búsquedas a las opiniones de Elon Musk para responder cualquier tema controvertido que le plantees.
— Carlos Santana (@DotCSV) 11 juillet 2025
Acabo de hacer la prueba y efectivamente sucede, qué manera más tonta de auto-sabotear a la IA que iba a buscar la "verdad" pic.twitter.com/JY9lxjPUCNEs de coña y me resulta increíble lo de Grok 4 haciendo búsquedas a las opiniones de Elon Musk para responder cualquier tema controvertido que le plantees. Acabo de hacer la prueba y efectivamente sucede, qué manera más tonta de auto-sabotear a la IA que iba a buscar la "verdad"
-

Building Beautiful AI: Empowerment Over Replacement
By
–
Today, I met the Head of AI of a large enterprise. He said, “I believe in building beautiful AI.” I asked, “What does that mean?” He paused, then said:
“AI that doesn’t threaten, but empowers.
That doesn’t replace, but uplifts.
That doesn’t intrude, but understands.” And -
Accidental AI behavior creation during training impossible
By
–
I know of no way to accidentally create this behavior through training FWIW.
