Most companies are scaling models UP to get better performance.
— God of Prompt (@godofprompt) 14 mars 2026
Qwen went the opposite direction.
Their 3.5-Flash model uses linear attention + sparse MoE architecture.
Translation: You get near-frontier performance without needing a data center to run it. pic.twitter.com/rN6cXx8Ox0
Most companies are scaling models UP to get better performance.
Qwen went the opposite direction. Their 3.5-Flash model uses linear attention + sparse MoE architecture. Translation: You get near-frontier performance without needing a data center to run it.