Here's a quick demo of the #Llama 3.2B chatbot running entirely on-device with Axelera AI’s Metis platform – quick to build, and easy to do! This 3B parameter model runs real-time at 7.1 tokens/sec per core — and thanks to our quad-core architecture, you can handle 4 chatbot
Llama 3.2B Chatbot Running On-Device with Metis Platform
By
–