Why is it that with ChatGPT, Gemini, Claude, Copilot and other LLMs we have to always start new chats for them to work well? What is the scientific explanation? What are the hypotheses? What is the evidence for each?
LLMS
-
Claude’s Post-Training Quality and Conversational Consistency Praised
By
–
Every time you are chatting with claude(Opus 4.5 esp. but seems to apply to all models), you quickly realize how clever claude post-training is. The attention to details, data quality. It's rare that claude will break in the middle of conversations. I think this is beyond system
-
DeepSeek Model Distillation Continues Strong Momentum
By
–
Looks like the DeepSeek distillation is still going strong
-
Claude’s Post-Training Excellence and Alignment in AI Models
By
–
Every time you are chatting with claude(Opus 4.5 esp. but seems to apply to all models), you quickly realize how clever claude post-training is. The attention to details, data quality. It's rare that claude will break in the middle of conversations. I think this is beyond system prompt and training dynamics/physics, and more about aura-like things: behaviors, personalities, when to push back/when to defer, and countless subtle edge cases. It's no surprise alignment people there follow closely the training. There was an invited talk at CMU this past spring on alignment and I asked why claude vibes and post-training feel different, the answer was high-level as you'd expect, but same: training, post-training, alignment, and evals teams are closed-loop. Sam Bowman (@sleepinyourhat) From everything we know so far, Opus 4.5 seems to be the best-aligned model out there in a bunch of ways. I follow the training process closely as part of my work on alignment evaluations. Here's my guess about the two things that are most responsible for making 4.5 special. 🧵 — https://nitter.net/sleepinyourhat/status/1997006353647522098#m
-

Poetiq beats Gemini 3 on ARC-AGI-2
By
–


Poetiq system achieved 54% score on ARC-AGI-2 benchmark, surpassing Gemini 3 Deep Think at more then twice lower compute cost. Poetiq (Mix) is a combined self learning system that leverages Gemini 3 and GPT-5.1 models. We want to do SVG tests
-
Claude Converts Speech to Rust Code with Whisper
By
–
Yeah: claude -p "Convert to Rust: $(whisper audio.mp3 -f txt)" | say
-

DeepSeek-V3.2: Open-Source Reasoning Model Rivals GPT-5
By
–
DeepSeek recently dropped a new open-source reasoning champion to rival GPT-5! Introducing DeepSeek-V3.2: 671B param Mixture-of-Experts (37B active) Gold-medal IMO & IOI performance Sparse attention for long contexts Balanced daily driver + tool-use Fully
-
NVIDIA Nemotron Models Integrated with Amazon Bedrock
By
–
NVIDIA Nemotron models are now integrated with Amazon Bedrock, making it easier to build and scale generative AI applications.
— NVIDIA AI (@NVIDIAAI) 6 décembre 2025
Early adopters are already deploying specialized agents with Nemotron on Bedrock:
🔒 @CrowdStrike is powering advanced security agents in Charlotte AI™… pic.twitter.com/lArJD8qO8SNVIDIA Nemotron models are now integrated with Amazon Bedrock, making it easier to build and scale generative AI applications. Early adopters are already deploying specialized agents with Nemotron on Bedrock: @CrowdStrike is powering advanced security agents in Charlotte AI™
-
The Blurring Line Between User and AI Prompting in 2026
By
–
The boundary between you prompting the model and the model prompting you is going to get blurry in 2026