Reinforcement Learning of Large Language Models, Spring 2025(UCLA) Great set of new lectures on reinforcement learning of LLMs. Covers a wide range of topics related to RLxLLMs such as basics/foundations, test-time compute, RLHF, and RL with verifiable rewards(RLVR).
GENERATIVE AI
-
AI Age Disruption: Product Market Fit Fragility Explained
By
–
Especially in the AI age. You could have PMF one day, and not the next.
-
AI is more than just language models
By
–
AI is a toolkit, but lately the hammer has been getting all the attention. So people say AI can’t cut, screw or drill.
-
OpenAI Using Claude Suggests It’s Genuinely Superior
By
–
even openai using claude means the alternative explanation is probably more true: claude is just that damn good and it isnt just mass hysteria
-
Edge Models Drive Adoption With Apache 2.0 License
By
–
Actually, edge models are (by far) the most demanded ones. The license applies to all derivatives, so it shouldn't be too confusing. We also modeled it after Apache 2.0 for user-friendliness.
-
New AI Model Too Similar to Claude Raises Concerns
By
–
And I would add (hopefully without creating too much speculation) it's a little *too* close to Claude, if you know what I mean.
-
Clarifying differences between Grok 4 Heavy and Grok 4
By
–
That’s not Grok 4 Heavy. The discrepancy between G4H and regular Grok 4 is discussed in my thread. I also included five Grok share links if you suspect my screenshots and screen recordings are fake.
-
Balancing Open-Source Contribution with Revenue Generation
By
–
It's possible in the future, but this license was created to balance our desire to contribute to the open-source community with the need to generate revenue. Unlike big tech companies, models are the only way we make money: no ads, no shady practices with user data. I hope
-
Harvey, Hebbia, and Major AI Benchmarks Performance
By
–
yeah… also harvey, hebbia, and all the benchmarks like MMLU, GPQA, HumanEval…
-
Anthropic Models Excel at Tool Calling in Loops
By
–
Overthinkers. I just need something that calls a tool, not a dissertation on why it's calling one. Which is why I like gpt-4.1 better, but even that one has flaws when in a loop. Basically, so far, only Anthropic models have been reliable for tools in a loop. This changes the
