确实很酷,拿 2900 万美金做 12M context,
— 艾略特 (@elliotchen100) 5 mai 2026
侧面证明了一件事:整个行业都开始相信稀疏注意力是 dense attention 的解药。
SubQ 走的是「重训一个模型」,属于垂直整合,风险大回报也大。@evermind 的 MSA 走的是「给主流模型加记忆」,属于 水平嵌入,谁的模型都能用。
另外,SubQ API 跟 SubQ… https://t.co/yZ8IugqquA
Yeah, that's pretty cool—using $29 million to build a 12M context window, which indirectly proves one thing: the entire industry is starting to believe that sparse attention is the antidote to dense attention. SubQ takes the "retrain a single model" approach, which is vertical