

One of the DeepSeek repositories got updated with a reference to a new “model1” model. “FlashMLA is DeepSeek's library of optimized attention kernels, powering the DeepSeek-V3 and DeepSeek-V3.2-Exp models.” Soon?
By
–



One of the DeepSeek repositories got updated with a reference to a new “model1” model. “FlashMLA is DeepSeek's library of optimized attention kernels, powering the DeepSeek-V3 and DeepSeek-V3.2-Exp models.” Soon?