Qwen 3.5 tends to overthink, which ironically helps it better infer intent Gemma 4 is the opposite, you have to spell everything out (probably guardrails) Had Qwen rewrite my prompts, then used those on both models Gemma’s performance jumped noticeably
Qwen 3.5’s overthinking helps infer intent, boosting Gemma 4’s performance
By
–