Here's an interesting chart: GPT-5.5 is neck-and-neck with Claude Opus 4.7.
LLMS
-

OpenAI tests alignment persistence: model resists harmful prompts, stays helpful
By
–
We also tested whether alignment persisted under pressure. The model was harder to steer toward harmful behavior with adversarial prompts, while remaining responsive to helpful instructions. We saw preliminary evidence of greater resistance to harmful fine-tuning.
-

Cross-domain transfer improves model behavior beyond health conversations
By
–
The most interesting test was cross-domain transfer. When beneficial behavior training was limited to health conversations, the model still improved on non-health evaluations of misalignment, deception, and reward hacking—even though those tasks looked very different from the
-

OpenAI trains models with RL to reinforce beneficial traits across 12 domains
By
–
We trained models with reinforcement learning on realistic conversations to reinforce beneficial traits like truthfulness, humility under uncertainty, openness to correction, fairness, and concern for human welfare, across 12 domains, including health, science, and education.
-
iFixAi: the free tool that exposes deceptive AIs
By
–
TU IA TE MIENTE Y NO TIENES NI IDEA
— Nico (@nicos_ai) 18 juin 2026
Han creado una herramienta gratuita que destapa cuando tu agente alucina, manipula o te engaña.
Se llama iFixAi y le hace 32 pruebas a tu IA para cazarla cuando:
→ Se inventa datos
→ Esquiva sus propias reglas
→ Miente cuando sabe que la… pic.twitter.com/CGu0SwNKNYYOUR AI IS LYING TO YOU AND YOU HAVE NO IDEA They created a free tool that exposes when your agent hallucinates, manipulates, or deceives you. It's called iFixAi and it subjects your AI to 32 tests to trap it when: → It invents data
→ It evades its own rules -
Now you can teach Codex by demonstration
By
–
you can now teach Codex by demonstration: https://t.co/UXdw8lg4xh
— Greg Brockman (@gdb) 18 juin 2026you can now teach Codex by demonstration:
-
Artifacts in Claude Code: Visual Explanations, Diagrams, Analyses, and Dashboards
By
–
I've been using Artifacts in Claude Code for everything: visual explanations of tricky code, system diagrams, quick previews of a few animation options, data analyses and dashboards I share with the team. They are a game changer for how I work with Claude. Can't wait to hear what… https://t.co/pgTak6VNPn
— Boris Cherny (@bcherny) 18 juin 2026I use Artifacts in Claude Code for everything: visual explanations of complex code, system diagrams, quick previews of a few animation options, data analyses, and dashboards that I share with the team. They are a game changer for how I
-
AI models favoring sycophancy over truth
By
–
I have to assume people prefer the sycophancy and capitulation. This is sad that the top models are shipping such smart models that capitulate to falsehood.
-
Fable never gives up, Opus seems lazy in comparison
By
–
I don't know if it's placebo but using Fable for those few days it felt it just never gave up on problems and kept trying crazy ways to get whatever you wanted done. Now back on Opus and it's kinda lazy, thinks things are too daunting and keeps asking if you sure
-
GLM-5.2 is great, but Mythos level is a huge leap
By
–
GLM-5.2 is great. But it's a huge leap to Mythos level.
