if you're not opting open source/ science – you're ngmi!
@reach_vb
-

MSFT Research Improves Mistral 7B with 25M Instruction Dataset
By
–
MSFT Research is Cracked AF! The full set 25M instruction set improves Mistral 7B – 40% on AGIEval, 19% on MMLU, 54% on GSM8K, 38% on BBH and 45% improvement on AlpacaEval ⚡
— Vaibhav (VB) Srivastav (@reach_vb) 15 novembre 2024
Please release the full set @MSFTResearch 🙏 https://t.co/zWpWlAvtZ7 pic.twitter.com/yVNrcXtaPDMSFT Research is Cracked AF! The full set 25M instruction set improves Mistral 7B – 40% on AGIEval, 19% on MMLU, 54% on GSM8K, 38% on BBH and 45% improvement on AlpacaEval Please release the full set @MSFTResearch
-
Microsoft Releases One Million Synthetic Instruction Pairs Dataset
By
–
Oh wow! @MSFTResearch released 1 MILLION synthetic instruction pairs covering different capabilities, such as text editing, creative writing, coding, reading comprehension, etc – permissively licensed 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 15 novembre 2024
Explore it directly on the Hugging Face Hub!
Kudos MSFT! Let the… pic.twitter.com/Y1BmQpi64zOh wow! @MSFTResearch released 1 MILLION synthetic instruction pairs covering different capabilities, such as text editing, creative writing, coding, reading comprehension, etc – permissively licensed Explore it directly on the Hugging Face Hub! Kudos MSFT! Let the
-
LibSQL: Open Source Database Fork Advances Developer Tools
By
–
https://
github.com/tursodatabase/
libsql
… -

Incredible developments emerging in the SQLite ecosystem
By
–
incredible things are happening in the sqlite world
-

Nexusflow Releases Athene v2 72B Competitive Model
By
–
Nexusflow released Athene v2 72B – competetive with GPT4o & Llama 3.1 405B Chat, Code and Math > Arena Hard: GPT4o (84.9) vs Athene v2 (77.9) vs L3.1 405B (69.3) > Bigcode-Bench Hard: GPT4o (30.8) vs Athene v2 (31.4) vs L3.1 405B (26.4) > MATH: GPT4o (76.6) vs Athene v2
-
SWE to MLE Transition: Enabling Engineers Career Paths
By
–
Congratulations on the legendary run! You’ve paved way for millions of SWEs transition to MLE and more. All the best for what’s next. Secretly hoping it’s Open Science/ Source too
-
Emilia Dataset Released for AI Model Training
By
–
Maybe this? https://
huggingface.co/datasets/amphi
on/Emilia-Dataset
… -
2 Trillion Multilingual High-Quality AI Training Tokens Released
By
–
2,003,039,184,047 multilingual, commercially permissive and high quality tokens!
