You got this. We believe in you. #DeepSeek
LLMS
-

DeepSeek R1: Rule-based RL beats SFT for GUI agents with smaller datasets
By
–
DeepSeek R1 moment has come for GUI agents: Rule-based Reinforcement Learning gives better results than SFT with 500x smaller datasets! Traditionally (by which I mean "in the last few months"), GUI agents have been trained with supervised fine-tuning (SFT). This meant,
-

Data Scaling Effects in Reinforcement Learning from Human Feedback
By
–
Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback Shen et al.: https://
arxiv.org/abs/2503.22230 #ArtificialIntelligence #DeepLearning #ReinforcementLearning -

Fastest AI Output Speed Verification: 267 Tokens Per Second
By
–
The GOAT of speed verifications @ArtificialAnlys Thank you for verifying our speeds! With up to 267 tokens/s, it's official — we're delivering the fastest output speeds out there!
-

LLM Text Interfaces Need to Evolve Beyond Command Terminals
By
–
Writing text back and forth with an LLM is like we're all the way back to the era of command terminals. The "correct" output is a lot closer to custom web apps written just for your query, information laid out spatially, multimodal, interactive, etc. Will take some time. https://t.co/dMAQEjE8GF pic.twitter.com/N849kmMTZQ
— Andrej Karpathy (@karpathy) 31 mars 2025Writing text back and forth with an LLM is like we're all the way back to the era of command terminals. The "correct" output is a lot closer to custom web apps written just for your query, information laid out spatially, multimodal, interactive, etc. Will take some time.
-

CPPO Accelerates Group Relative Policy Optimization Training
By
–
CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning Models
Paper: https://
arxiv.org/pdf/2503.22342
Code: https://
github.com/lzhxmu/CPPO -

CPPO Boosts GRPO Speed by 8x on Mathematical Reasoning
By
–
GRPO just got a speed boost! Xiamen University introduced Completion Pruning Policy Optimization (CPPO), which significantly reduces the number of gradient calculations and updates.
How fast? On GSM8K, it's 8.32× faster than GRPO, and on MATH, the speedup is 3.51×. -

LLaMA 4 and New LLM Models Spider Cybele Moonhowler Emerge
By
–
LLaMA 4 is coming?? New models named “Spider” and “Cybele" have appeared in the LMSYS Arena. Also, a model that appears to be from Google, called Moonhowler, has shown up as well.
-

Baidu’s ERNIE 4.5 defeats OpenAI’s GPT-4.5 at Chinese chess
By
–
Baidu's ERNIE 4.5 model just destroyed OpenAI's GPT-4.5 at Chinese chess ERNIE won all three matches and was even seen “taking it easy” during portions of the lopsided victories!
