TinyZero reproduction of R1-Zero
"experience the Ahah moment yourself for < $30" Given a base model, the RL finetuning can be relatively very cheap and quite accessible.
OPEN SOURCE
-

TinyZero: Affordable RL Finetuning Under $30
By
–
-
Building Diverse RL Environments for LLM Cognitive Strategy Development
By
–
For friends of open source: imo the highest leverage thing you can do is help construct a high diversity of RL environments that help elicit LLM cognitive strategies. To build a gym of sorts. This is a highly parallelizable task, which favors a large community of collaborators.
-
DeepSeek R1 Sonar for search optimization
By
–
DeepSeek R1 powered Sonar implementation (optimised for search use-cases)
-
Grok 3 to support reasoning and UI exposure
By
–
BREAKING 🚨: Grok 3 will support reasoning! It will be able to expose its "thinking" process to the UI as well 👀 https://t.co/ugfgzc2So3 pic.twitter.com/nkL0Jx8Djv
— 🚨 AI News | TestingCatalog (@testingcatalog) 29 janvier 2025BREAKING : Grok 3 will support reasoning! It will be able to expose its "thinking" process to the UI as well
-
LLM Distillation Term Usage: R1 Dataset Curation and Model Training
By
–
Today, in LLM contexts the term "distillation" is used quite loosely. In the case of R1 it just means that they created and curated a dataset for SFT from R1 that they used to train distilled R1 models based on Qwen and Llama.
-

LangGraph State Management Primitives Now Accessible Beyond Framework
By
–
excited to make low level primitives for state management accessible outside of LangGraph next up: higher level, off-the-shelf implementations for short and long term memory
-
Custom Queue Integration with Drupal 11 Framework
By
–
Learn how to create a custom queue, along with the queue factory needed to integrate that queue, with @Drupal 11 https://
bit.ly/3CrQ4M1 #drupal #drupal11 #opensource #openweb -
DeepSeek Qwen 7B Abliterated Released Next Version
By
–
deepseek qwen 7b abliterated has been submitted w/ the next release
-
Open Source Requirements: Data Transparency Beyond Code and Weights
By
–
For something to be open source, we need to see 1. Data it was trained & evaluated on
2. Code
3. Model architecture
4. Model weights. DeepSeek only gives 3, 4. And I'll see the day that anyone gives us #1 without being forced to do so, because all of them are stealing data. -
Running R1 Distillations on Qwen and Llama Models
By
–
you can run the distillations of r1 which are tunes on qwen and llama – full r1 needs a few hundred gb of ram