🔥 𝐑𝐞𝐥𝐞𝐚𝐬𝐢𝐧𝐠 𝐈𝐧𝐝𝐢𝐜 𝐆𝐞𝐦𝐦𝐚 7𝐁/2𝐁 𝐈𝐧𝐬𝐭𝐫𝐮𝐜𝐭𝐢𝐨𝐧 𝐭𝐮𝐧𝐞𝐝 𝐦𝐨𝐝𝐞𝐥 𝐨𝐧 9 𝐈𝐧𝐝𝐢𝐚𝐧 𝐋𝐚𝐧𝐠𝐮𝐚𝐠𝐞𝐬 — 𝐍𝐚𝐯𝐚𝐫𝐚𝐬𝐚 🚀 We are thrilled to share 🌟 𝐍𝐚𝐯𝐚𝐫𝐚𝐬𝐚, a Gemma 7B & 2B instruction-tuned models in 9 Indian Languages – Perhaps this is the first Indic open instruction-tuned model trained in 9 Indian languages additionally English included. 🔥𝐍𝐚𝐯𝐚𝐫𝐚𝐬𝐚 is a Gemma 7B & 2B SFT model using Gemma 7B & 2B base models. Last week we released the Telugu Gemma 7B/ 2B SFT model using curated Telugu datasets from Telugu LLM Labs and we observed really good performance compared to Llama2-based models. 🌐 So, we thought why don’t we scale up Gemma 7B & 2B models to multiple Indian languages and we went ahead with testing tokenizers of the following 9 Indian Languages and English Language. 1. Hindi 2. Telugu 3. Tamil 4. Malayalam 5. Kannada 6. Gujarati 7. Bengali 8. Punjabi 9. Odia 10. English ✨ We found the model to have the following capabilities: (X represents any other Indian language) 1. Instruction and Input in Native X language, Output in Native X language. 2. Instruction and Input in English language prompted to respond in Native X language, Output in Native X language. 3. Instruction in Native X language, Input in English language, and Output in Native X language. 📊𝐓𝐫𝐚𝐢𝐧𝐢𝐧𝐠 𝐃𝐞𝐭𝐚𝐢𝐥𝐬: 1. Single A100 machine which took approx. 36 hours for the 7B model and 15 hours for the 2B model. 2. Platform: E2E Networks Limited 📝 We have shared details on datasets, Examples of Reasoning, Translation, and Question Answering with Context in our blog post. 🤝 The work would not have been possible without huge community effort from different languages and a huge shout out to each one of their work over the past few months showcasing the true OSS power. Following are details of contributors for the languages: 1. Hindi: @SarvamAI 2. Telugu: Telugu LLM Labs 3. Tamil: @abhinand58 4. Kannada: @adarshxs and the team at Tensonic 5. Malayalam: Vishnu Prasad J 6. Odia: @OdiaGenAI 7. Gujarati: Adarsh Shirawalmath and the team at Tensonic 8. Punjabi: HydraIndicLM 9. Bengali: HydraIndicLM 👏 Special thanks to @unslothai for simplifying the training and inference processes! 🔜 As we release these models, the next step is to create romanized datasets and we are working hard on evaluation datasets so that we can benchmark and improve on top of it. 🤝 This work is done in collaboration with @ramsri_goutham as part of the Telugu LLM Labs independent initiative. 𝐁𝐥𝐨𝐠𝐏𝐨𝐬𝐭: shorturl.at/jBQWY 𝐂𝐨𝐝𝐞𝐁𝐚𝐬𝐞: shorturl.at/elxBF
→ View original post on X — @sudalairajkumar, 2024-03-06 05:14 UTC