AI Dynamics

Global AI News Aggregator

About

Real Chat Tests on Four AI Models

the experiment is clean they took real multi-turn conversations from WildChat and ShareLM. not synthetic benchmarks. actual human-ai chats then they ran every conversation two ways across four models (Qwen3-4B, DeepSeek-R1-8B, GPT-OSS-20B, and GPT-5.2): > full context: normal.

→ View original post on X — @godofprompt