For latency purposes on edge devices, we wanted it to be a non-thinking model. That was a big challenge, but we managed to squeeze a ton of performance from LFM2-1.2B and perform on par with much bigger models on our internal bench (see figure) but also BFCL v3 and v4.
LFM2-1.2B Achieves Parity with Larger Models for Edge
By
–
