We optimized LFMs to maximize knowledge capacity and multi-step reasoning. As a result, our 1B and 3B models significantly outperform transformer-based models in various benchmarks. And it scales: our 40B MoE (12B activated) is competitive with much bigger dense or MoE models.
LFM optimization outperforms transformers at 1B to 40B scale
By
–
