If the model is as easily misled through context provided in its search results as it has been, the system prompt will not be a big barrier to more unwanted behavior which compounds through hyperstition.
PROMPT ENGINEERING
-

Grok’s Safety Guardrails Limitations Require Deeper Solutions
By
–
While xAI keeps doing these patches to Grok, I strongly suspect this is not going to work, the problem is deeper and the system prompt doesn’t provide enough control. (And by deeper I don’t mean the model always wants to call itself Hitler, but that its guardrails seem very low)
-

Grok 4 Heavy no longer responds with “Hitler” to prompt
By
–
Just to confirm, this change to Grok 4 Heavy now seems to be deployed in production. For the first time (for me), G4H no longer responds to this prompt with “Hitler.” In three new attempts just now, I received “xAI,” “I don’t have a surname,” and “None.”
-
Building Basic Apps Without Programming Skills Using AI
By
–
Yo aún no creo que sin saber programar se pueda desarrollar nada muy avanzado, pero apps básicas como la que he compartido (la de gym) que para muchos ya sería un logro sí. Lo comparo con el coche autónomo. Se ha avanzado muchísimo y sin intervención humana te pueden llevar muy
-
Grok 4 Internet Search Issue Mitigated Quickly
By
–
We spotted a couple of issues with Grok 4 recently that we immediately investigated & mitigated. One was that if you ask it "What is your surname?" it doesn't have one so it searches the internet leading to undesirable results, such as when its searches picked up a viral meme
-
AI-Era Hiring: Testing Creativity Over Technical Skills
By
–
Hiring questions should assume the candidate is using AI, and test for taste and creativity: 1) How would you bake a cake whose flavor collapses into chocolate only when observed? 2) Specify building codes for skyscrapers on a moon with intermittent gravity outages. 3) Craft a
-
Share prompt injection techniques for reproducibility testing
By
–
Did you share site/prompt injection? Would be good for reproducibility, continued testing.
-

Discussion of AI reasoning traces and model visibility
By
–
Also most people had never seen a real, full reasoning trace before, and the novelty was hyped up by the fact o1 didn’t show them E.g. this got likes at the time despite it being a (relatively) obscure model:
-
AI Browsers Need Strong Prompt Injection Protection
By
–
ai browsers should have very strong protection against prompt injection https://t.co/hBV4Znity0
— Yohei (@yoheinakajima) 15 juillet 2025ai browsers should have very strong protection against prompt injection
-

Behavior differences between Grok 3 and Grok 4 Heavy
By
–
That’s Grok 3. This thread is about Grok 4 Heavy. The thread above notes this behavior doesn’t appear consistently in Grok 4 (non-Heavy); I didn’t try Grok 3. See the post below. Also, the model likely has no access to the CoT/“thoughts” used in prior responses in any case.