AI Dynamics

Global AI News Aggregator

About

Navarasa: Gemma 7B/2B Instruction-Tuned Model for 9 Indian Languages

๐Ÿ”ฅ ๐‘๐ž๐ฅ๐ž๐š๐ฌ๐ข๐ง๐  ๐ˆ๐ง๐๐ข๐œ ๐†๐ž๐ฆ๐ฆ๐š 7๐/2๐ ๐ˆ๐ง๐ฌ๐ญ๐ซ๐ฎ๐œ๐ญ๐ข๐จ๐ง ๐ญ๐ฎ๐ง๐ž๐ ๐ฆ๐จ๐๐ž๐ฅ ๐จ๐ง 9 ๐ˆ๐ง๐๐ข๐š๐ง ๐‹๐š๐ง๐ ๐ฎ๐š๐ ๐ž๐ฌ โ€” ๐๐š๐ฏ๐š๐ซ๐š๐ฌ๐š ๐Ÿš€ We are thrilled to share ๐ŸŒŸ ๐๐š๐ฏ๐š๐ซ๐š๐ฌ๐š, a Gemma 7B & 2B instruction-tuned models in 9 Indian Languages – Perhaps this is the first Indic open instruction-tuned model trained in 9 Indian languages additionally English included. ๐Ÿ”ฅ๐๐š๐ฏ๐š๐ซ๐š๐ฌ๐š is a Gemma 7B & 2B SFT model using Gemma 7B & 2B base models. Last week we released the Telugu Gemma 7B/ 2B SFT model using curated Telugu datasets from Telugu LLM Labs and we observed really good performance compared to Llama2-based models. ๐ŸŒ So, we thought why donโ€™t we scale up Gemma 7B & 2B models to multiple Indian languages and we went ahead with testing tokenizers of the following 9 Indian Languages and English Language. 1. Hindi 2. Telugu 3. Tamil 4. Malayalam 5. Kannada 6. Gujarati 7. Bengali 8. Punjabi 9. Odia 10. English โœจ We found the model to have the following capabilities: (X represents any other Indian language) 1. Instruction and Input in Native X language, Output in Native X language. 2. Instruction and Input in English language prompted to respond in Native X language, Output in Native X language. 3. Instruction in Native X language, Input in English language, and Output in Native X language. ๐Ÿ“Š๐“๐ซ๐š๐ข๐ง๐ข๐ง๐  ๐ƒ๐ž๐ญ๐š๐ข๐ฅ๐ฌ: 1. Single A100 machine which took approx. 36 hours for the 7B model and 15 hours for the 2B model. 2. Platform: E2E Networks Limited ๐Ÿ“ We have shared details on datasets, Examples of Reasoning, Translation, and Question Answering with Context in our blog post. ๐Ÿค The work would not have been possible without huge community effort from different languages and a huge shout out to each one of their work over the past few months showcasing the true OSS power. Following are details of contributors for the languages: 1. Hindi: @SarvamAI 2. Telugu: Telugu LLM Labs 3. Tamil: @abhinand58 4. Kannada: @adarshxs and the team at Tensonic 5. Malayalam: Vishnu Prasad J 6. Odia: @OdiaGenAI 7. Gujarati: Adarsh Shirawalmath and the team at Tensonic 8. Punjabi: HydraIndicLM 9. Bengali: HydraIndicLM ๐Ÿ‘ Special thanks toย @unslothai for simplifying the training and inference processes! ๐Ÿ”œ As we release these models, the next step is to create romanized datasets and we are working hard on evaluation datasets so that we can benchmark and improve on top of it. ๐Ÿค This work is done in collaboration withย @ramsri_gouthamย as part ofย the Telugu LLM Labsย independent initiative. ๐๐ฅ๐จ๐ ๐๐จ๐ฌ๐ญ: shorturl.at/jBQWY ๐‚๐จ๐๐ž๐๐š๐ฌ๐ž: shorturl.at/elxBF

โ†’ View original post on X โ€” @sudalairajkumar, 2024-03-06 05:14 UTC