[1] https://www.theverge com/24066646/ai-electricity-energy-watts-generative-consumption [2] https://www.numenta com/blog/2023/08/10/ai-is-harming-our-planet-2023/
GENERATIVE AI
-

GPT-4 Training Energy Consumption Versus Netflix Streaming Hours
By
–
"it’s estimated that GPT-4 consumed between 51,773 MWh and 62,319 MWh" [1] "streaming an hour of Netflix requires around 0.8 kWh (0.0008 MWh) of electricity." [2] Es decir, entrenar a GPT-4 (tomando el valor superior) cuesta unas ~78M horas de Netflix. Es decir, el costo
-
OpenAI Cofounder Departure to Anthropic Strengthens Competition
By
–
¿Cómo se debe de sentir internamente cuando uno de los cofundadores de OpenAI se marcha de la empresa para irse a Anthropic? Personalmente me encanta que Anthropic esté ganando fuerza porque lo están haciendo muy bien. Y la competencia siempre nos beneficia 🙂
-

Mistral Large 2 Excels in Coding, Math, and Hard Prompts
By
–
Mistral Large 2 (2407) is now on @lmsysorg
. It performs extremely well in the Coding, Hard Prompts, Math, and Longer Query categories, where it outperforms GPT4-Turbo and Claude 3 Opus. It is also doing very well in Instruction Following where it ranks above Llama 3.1 405B. -
Whisper Model Customization for Air Traffic Control Safety
By
–
Enhance Speech Recognition in Air Traffic Control with Whisper. We've tailored OpenAI's powerful Whisper model to better understand the unique challenges of air traffic communications.
— Satya Mallick (@LearnOpenCV) 6 août 2024
Discover how customizing Whisper can improve safety and accuracy, with insights into our… pic.twitter.com/pMERB815ziEnhance Speech Recognition in Air Traffic Control with Whisper. We've tailored OpenAI's powerful Whisper model to better understand the unique challenges of air traffic communications. Discover how customizing Whisper can improve safety and accuracy, with insights into our
-
Open-source AI models: transparency over performance gains
By
–
Open-source AI models are like sport without doping Performances can be slightly lower but in the end, transparency is more sustainable, healthy and exciting to watch and participate in 🙂
-
Experimenting with Model Merging and Quantization Techniques
By
–
Anyway, have fun quantizing and running this monster! This is a silly experiment but who knows, we might: 1/ Realize it's good at something, like creative writing
2/ Get some insights into how to create these self-merges Good luck everyone and make your own frankenmerges -

BigLlama-3.1-681B-Instruct: New Intermediate Checkpoint Released
By
–
Because 1T parameters might be a bit too much, I'm also releasing an "intermediate checkpoint": BigLlama-3.1-681B-Instruct. It's closer to what Llama 3 120B was to the 70B, so it might be more coherent as well. Model: https://
huggingface.co/mlabonne/BigLl
ama-3.1-681B-Instruct
… -

BigLlama-3.1 reaches 1 trillion parameters milestone
By
–
BigLlama-3.1-1T-Instruct So I've heard that 405B parameters weren't enough… It's my pleasure to present an upscaled Llama 3.1 with 1,000,000,000 parameters. Now available on @huggingface
. Model: https://
huggingface.co/mlabonne/BigLl
ama-3.1-1T-Instruct
… -
OpenAI GPT-6 training scale and deployment timeline
By
–
Esto es como cuando te tomas un break para el café mientras se entrena tu modelo, pero claro… a la escala de OpenAI. Lanzas el entrenamiento de GPT-6 y te tomas unos meses sabáticos hasta que termine.