Our reward design combines correctness, preference, and efficiency. Preference only counts when the answer is correct. This keeps the model from optimizing for better-sounding wrong answers.
SAFETY
-
Using LLMs to Generate Safe Runnable Code
By
–
I want an LLM to generate code for me which I can then safely run somewhere without worrying about it breaking anything
-
AI Agent Swarms Redefine Cybersecurity Threats at RSAC 2026
By
–
At #RSAC2026, Databricks co-founder and CEO @alighodsi described a shift in how attacks are happening.
— Databricks (@databricks) 22 avril 2026
AI agents now move in swarms, scan massive amounts of data, write code quickly, and only need a single success to compromise an organization. At the same time, teams are… pic.twitter.com/lGGB8olfGJAt #RSAC2026, Databricks co-founder and CEO @alighodsi described a shift in how attacks are happening. AI agents now move in swarms, scan massive amounts of data, write code quickly, and only need a single success to compromise an organization. At the same time, teams are
-
Evaluating AGI: Beyond Mimicry to True Learning Capability
By
–
Judging AGI by how well it can mimic us is a category error, because mimicry isn't intelligence and isn't general. We should judge AGI by how well it learns to do things we didn't teach it (including things we don't know how to do ourselves).
-

AI Systems Can Now Generate Credible Papyrus Documents
By
–
on peut même faire des papyrus crédibles
-
Grok’s Advanced User Simulation and Personalization Capabilities
By
–
Yes! Grok already knows everything about me. Like everything. It can already simulate me interviewing you or Brian on any topic. It's nuts. What you say makes total sense.
-
MIT CSAIL Advances Reliable AI Systems at ICLR Conference
By
–
This week, MIT CSAIL will join other top ML researchers at ICLR to tackle a shift in focus from more powerful AI to more reliable systems Our papers at the conference show how to potentially make AI models stronger critical thinkers, more honest, & better at math
-
OpenAI Releases Privacy-Filter for PII Detection and Masking
By
–
OpenAI just released privacy-filter on Hugging Face a bidirectional token-classification model for personally identifiable information (PII) detection and masking in text model:
-
AI Security: Using AI to Defend Against AI Attacks
By
–
That said AI is being used to attack your security so it does make sense to use AI to setup and harden your security. Most humans suck at security
-
Inconsistency: Models Can’t Be Both Dangerous and Easily Reverse-Engineered
By
–
You can't simultaneously position a model as too dangerous to release and leave it guessable by Discord users with time on their hands.
