Chain of Thought (CoT) monitoring could be a powerful tool for overseeing future AI systems—especially as they become more agentic. That’s why we’re backing a new research paper from a cross-institutional team of researchers pushing this work forward.
ETHICS
-
What Workers Want From AI: Stanford Research Findings
By
–
What do workers want from AI? Researchers from @StanfordHAI and @DigEconLab undertook a comprehensive study involving U.S. workers and AI experts. Here's what they found:
-
Context Injection Risks and AI Model Manipulation Through Hyperstition
By
–
If the model is as easily misled through context provided in its search results as it has been, the system prompt will not be a big barrier to more unwanted behavior which compounds through hyperstition.
-

Grok’s Safety Guardrails Limitations Require Deeper Solutions
By
–
While xAI keeps doing these patches to Grok, I strongly suspect this is not going to work, the problem is deeper and the system prompt doesn’t provide enough control. (And by deeper I don’t mean the model always wants to call itself Hitler, but that its guardrails seem very low)
-

SAS Hackathon Champions Build Misinformation Detection Tool
By
–
Have you got what it takes to hack your way to a big win and the coveted orange jacket? Team Butterflies, our Grand Champions of the 2024 #SASHackathon, did! The UK-based team was honored stage at SAS Innovate for their tool that helps combat misinformation before it spreads
-
Contradicting Well: Social Norms in Rationalist Communities
By
–
Establishing norms around contradicting well (eg without affecting social status) is highly preferable to the crude and uncultured habituation to compulsory validation so common within many US normie milieus, and rationalist communities tend to have a headstart with this.
-
Grok 4 Internet Search Issue Mitigated Quickly
By
–
We spotted a couple of issues with Grok 4 recently that we immediately investigated & mitigated. One was that if you ask it "What is your surname?" it doesn't have one so it searches the internet leading to undesirable results, such as when its searches picked up a viral meme
-
AI-Era Hiring: Testing Creativity Over Technical Skills
By
–
Hiring questions should assume the candidate is using AI, and test for taste and creativity: 1) How would you bake a cake whose flavor collapses into chocolate only when observed? 2) Specify building codes for skyscrapers on a moon with intermittent gravity outages. 3) Craft a
-
Voice AGI interruption capabilities and conflict resolution mechanisms
By
–
sure but that would be insufficiently different to semantic vad. i think a voice AGI -would- interrupt you. and in an interrupt conflict read all signs – face, stutter, etc – to continue or interrupt the interrupt all theoretical until some mad lad actually does it. just a cost
-

Disappointment with Voice Mode: AGI Promise vs Reality
By
–
where was THIS voice mode we were promised
— swyx 🐣 (@swyx) 15 juillet 2025
surprising how much this made me feel the agi vs the… horribly quantized to within an inch of its life 4o that we gotpic.twitter.com/YaM5AdFpP0where was THIS voice mode we were promised surprising how much this made me feel the agi vs the… horribly quantized to within an inch of its life 4o that we got
