3). DeepSeek-V2 – a strong MoE 236B parameter model, of which 21B are activated for each token; supports a context length of 128K tokens and uses Multi-head Latent Attention (MLA) for efficient inference by compressing the Key-Value (KV) cache into a latent vector…
DeepSeek-V2: 236B MoE Model with Efficient Latent Attention
By
–