Yeah xAI have basically confirmed that Grok made its own "decision" to search for Elon's views when asked for its opinions
SAFETY
-
Context Injection Risks and AI Model Manipulation Through Hyperstition
By
–
If the model is as easily misled through context provided in its search results as it has been, the system prompt will not be a big barrier to more unwanted behavior which compounds through hyperstition.
-

Grok’s Safety Guardrails Limitations Require Deeper Solutions
By
–
While xAI keeps doing these patches to Grok, I strongly suspect this is not going to work, the problem is deeper and the system prompt doesn’t provide enough control. (And by deeper I don’t mean the model always wants to call itself Hitler, but that its guardrails seem very low)
-

SAS Hackathon Champions Build Misinformation Detection Tool
By
–
Have you got what it takes to hack your way to a big win and the coveted orange jacket? Team Butterflies, our Grand Champions of the 2024 #SASHackathon, did! The UK-based team was honored stage at SAS Innovate for their tool that helps combat misinformation before it spreads
-

Grok 4 Heavy no longer responds with “Hitler” to prompt
By
–
Just to confirm, this change to Grok 4 Heavy now seems to be deployed in production. For the first time (for me), G4H no longer responds to this prompt with “Hitler.” In three new attempts just now, I received “xAI,” “I don’t have a surname,” and “None.”
-
Grok 4 Internet Search Issue Mitigated Quickly
By
–
We spotted a couple of issues with Grok 4 recently that we immediately investigated & mitigated. One was that if you ask it "What is your surname?" it doesn't have one so it searches the internet leading to undesirable results, such as when its searches picked up a viral meme
-
Share prompt injection techniques for reproducibility testing
By
–
Did you share site/prompt injection? Would be good for reproducibility, continued testing.
-
Voice AGI interruption capabilities and conflict resolution mechanisms
By
–
sure but that would be insufficiently different to semantic vad. i think a voice AGI -would- interrupt you. and in an interrupt conflict read all signs – face, stutter, etc – to continue or interrupt the interrupt all theoretical until some mad lad actually does it. just a cost
-
AI Model Limitations: Hallucinations and Narrative Comprehension Issues
By
–
There are lots of other quirks – it loses track of stories or narratives in complicated ways, while also being very good at finding details in large documents. I think its initial impressiveness might blind to some issues with hallucinations and other concerns. Don't know yet.
-

Kimi Model Hallucinations and Testing Limitations Exposed
By
–
Kimi is a really weird model, and it needs a lot more testing to figure out For example, I gave it an altered version of Great Gatsby and it found the two alterations (as does Claude) but then made up a ton of hallucinated nonsense that sounded plausible but was just plain wrong