"Reinforcement Learning: An Introduction" Richard S. Sutton and Andrew G. Barto FULL PDF: http://
incompleteideas.net/book/RLbook202
0.pdf
… #Artificialintelligence #DeepLearning #ReinforcementLearning
OPEN SOURCE
-

Reinforcement Learning Introduction by Sutton and Barto PDF
By
–
-

DeepSeek Releases Janus-Pro-7B Open-Source Multimodal AI Model
By
–
DeepSeek just dropped ANOTHER open-source AI model, Janus-Pro-7B. It's multimodal (can generate images) and beats OpenAI's DALL-E 3 and Stable Diffusion across GenEval and DPG-Bench benchmarks. This is just super cool.
Now they will move their side project as their main project -
Qwen 2.5 7B Released with 1M Context Window
By
–
Qwen 2.5 7B just dropped with a 1M context window. Here it is running at 4bit @ 11 tok/s.
— Aaron Ng (@localghost) 27 janvier 2025
Should we add it to the next update? 1M of context adds a lot of possibilities. pic.twitter.com/y4uhLfHKNSQwen 2.5 7B just dropped with a 1M context window. Here it is running at 4bit @ 11 tok/s. Should we add it to the next update? 1M of context adds a lot of possibilities.
-
DeepSeek R1 True Training Costs Beyond $5.567M
By
–
The R1 model wasn't free, and the DeepSeek R1 paper doesn't talk about costs at all. The $5.567M cost for DeepSeek is for when you have all the ingredients in front of you – running that recipe costs $5.567M. It doesn't consider the cost of procuring those ingredients (in this
-
llama.cpp enables easy local model execution
By
–
llama.cpp is the best! It’s wild that one can run a model of this size at the click of a button
-
Adding Official OpenAI Transformer Weights to Notebook
By
–
Ah nice! I thought they only had the transformer-library-formatted weights on the hub! In that case, I can actually add that to the notebook directly. Cool stuff! (Btw the reason why I opted for the OpenAI Tf ones was that they were the original/official ones)
-
Lower-precision formats and distilled models enabling local long-context LLMs
By
–
Me neither :). I think with new lower-precision formats, and better distilled models, we'll maybe increasingly adopt long-context LLMs run locally.
-

DeepSeek-R1 Training Cost: Missing R1 Distillation Expenses
By
–
The real question: What is the DeepSeek-R1 training cost? The $5.567M DeepSeek cost is missing the cost of training the R1 model to get distilled data. Table 9 in the DeepSeek v3 paper shows that the R1 distillation step is critical for quality. The R1 paper doesn't talk about
-
Loading weights from Hugging Face model hub as alternative
By
–
Sorry about the hassle loading the weights directly. But I actually have some bonus material on that to load them from the HF model hub as an alternative (should be referenced in the book itself):
-

Chinese Company Releases Open-Source Multimodal AI Model Janus
By
–
Le temps que les gens découvrent Deepseek R-1, l'entreprise chinoise a sorti, il y a une heure, un modèle open-source multimodal, "Janus", de génération et de compréhension d'images. (Vous voyez où on en est en termes de vélocité du marché sur l'IA ?) Lien ci-dessous.