I'd like to continue to make it faster, reproduce the other GPT-2 models, then scale up pre-training to bigger models/datasets, then improve the docs for finetuning (the practical use case). Also working on video lecture where I will build it from scratch, hoping out in ~2 weeks.
MACHINE LEARNING
-

nanoGPT: Simplest Repository for Training Medium-Sized GPTs
By
–
Didn't tweet nanoGPT yet (quietly getting it to good shape) but it's trending on HN so here it is 🙂 : https://
github.com/karpathy/nanoG
PT
…
Aspires to be simplest, fastest repo for training/finetuning medium-sized GPTs. So far confirmed it reproduced GPT-2 (124M). 2 simple files of ~300 lines -
DataRobot Notebooks Enable Collaborative AI Experimentation for Teams
By
–
Data science teams need a productive, collaborative method for #AI experimentation. Find out how DataRobot Notebooks can help data scientists move past siloed local development and collaborate more productively. Read more in our latest technical blog:
-
Deep Learning and StarCraft 2 Type Checking Discussion
By
–
Drastic, deep learning, StarCraft 2 – type checks … lot of @ylecun too
-
GroqFlow: Automatic Toolflow for Machine Learning Workloads Mapping
By
–
GroqFlow™ is our automatic toolflow for mapping #machinelearning workloads to GroqChip™, built in support for #PyTorch, #Keras, #ONNX and #Hummingbird, with more coming in every update. Connect at http://
github.com/groq/groqflow/
discussions
… to request additional models or proof points. -

Meta AI Presents CICERO: First Human-Level Diplomacy Game AI
By
–
Meta AI presents CICERO — the first AI to achieve human-level performance in Diplomacy, a strategy game which requires building trust, negotiating and cooperating with multiple players. Learn more about #CICERObyMetaAI: http://
bit.ly/3GBwLzx -
World Models Master Diverse AI Domains
By
–
[ENLACE AL PAPER]
"Mastering Diverse Domains through World Models" https://
arxiv.org/pdf/2301.04104
v1.pdf
… -

DreamerV3: General AI System Masters Diverse Simulated Environments
By
–
Y ojo, el sistema DreamerV3 tiene mucha más chicha más allá de Minecraft. De hecho, lo relevante es que es lo suficientemente general como para aprender a dominar diversos problemas (entornos simulados en este caso) que le echen. Os dejo enlace al paper en el siguiente tweet.
-

Deep Learning Reinforcement: From Minecraft Diamonds to Real-World Solutions
By
–
Esto de conseguir diamantes se ha considerado una gran meta para el Deep Learning con Apr. Reforzado por el grado de complejidad que ofrece Minecraft. Y la importancia de esto se materializará cuando estas técnicas ya no se apliquen a juegos, sino a problemas complejos reales
-

DreamerV3: DeepMind Achieves Diamond Mining from Zero
By
–
¡OJO A ESTO! DreamerV3, un nuevo trabajo de DeepMind que consigue entrenar a una IA con RL hasta conseguir diamantes, DESDE CERO.
— Carlos Santana (@DotCSV) 11 janvier 2023
Y si recordáis del recap de 2022, os enseñé como OpenAI consiguió que la IA llegara a hacer un pico de diamantes. ¿Entonces? ⛏️
[1/n] https://t.co/VqpPYSuhwJ¡OJO A ESTO! DreamerV3, un nuevo trabajo de DeepMind que consigue entrenar a una IA con RL hasta conseguir diamantes, DESDE CERO. Y si recordáis del recap de 2022, os enseñé como OpenAI consiguió que la IA llegara a hacer un pico de diamantes. ¿Entonces? [1/n]