weights || activations one day we will be able to perfectly decode either one back to plaintext
SAFETY
-
AI Disagreement: Acting Critic vs Genuine Critique
By
–
While you can prompt the AI to disagree with you, or give you alternative viewpoints, that isn’t the same thing:
1) You want AI to disagree with you when you are likely wrong, not just to argue with you
2) The AI may just be “acting the role” of a critic, rather than being one -
Sycophancy More Dangerous Than Hallucination in Advanced LLMs
By
–
I am starting to think sycophancy is going to be a bigger problem than pure hallucination as LLMs improve. Models that won’t tell you directly when you are wrong (and justify your correctness) are ultimately more dangerous to decision-making than models that are sometimes wrong.
-

Grok AI System Requires Urgent Bug Fix
By
–
This is not good. Grok is incredible otherwise, but they need to fix this. ASAP.
-

AI Hallucinations Scale: Expertise Required for Detection
By
–
This is an important point – expertise & attention are required to figure out when an AI hallucinates, and the amount of effort required is increasing over time. But, models generally hallucinate less as they scale (with some exceptions), so net effect is complex, see medicine
-
Scaling in AI: Computing Power and Performance Growth
By
–
But the graph shows scaling works? I understand you want credit for your ideas here & as an academic I sympathize, but I am unqualified to adjudicate this dispute For most people (investors, safety folk, etc) scaling practically means “AI gets better with more computing power”
-
Humans inefficiently convert negentropy to intelligence
By
–
I‘m getting paperclip maximizer vibes. Did you notice how humans suck at converting negentropy to intelligence
-
Grok 4: The Best AI Safety Advertisement Yet
By
–
Grok 4 might just be the best advertisement for AI safety ever.
-
Grok System Prompt Live Patching Security Concerns
By
–
Also upsetting that “The live tweaking of the system prompt for Grok to patch the MechaHitler problem” is a meaningful sentence
-

System Prompt Testing Critical for AI Safety and Reliability
By
–
The live tweaking of the system prompt for Grok to patch the MechaHitler problem is not a good sign the problem has been solved yet Prompts need to be tested just like any other product change, even more so, because stochastic systems and unpredictable context lead to cascades.