Llama 4 is our first collection of models built using a mixture of experts (MoE) architecture. This architecture is more compute efficient for model training and inference and delivers higher quality models compared to dense architectures.
Llama 4: Mixture of Experts Architecture for Efficient Models
By
–
