Today we’re releasing TRL v1. 75+ methods. SFT, DPO, GRPO, async RL to take advantage of the latest and greatest open-source. 6 years from first commit to the library that post-trains most open models in the world. Built to be future proof. pip install trl
LLMS
-

Pareto Chart Analysis for Local LLMs Performance
By
–
this is the right Pareto chart for local LLMs, everything else is just pretending (they should use AAIIv4 instead of doing their own average but eh, nitpicks)
-

Qwen3.5-Omni Plus and Flash Models Now Available on Poe
By
–
Qwen3.5-Omni Plus and Qwen3.5-Omni Flash are now live on Poe. Both models understand text, images, audio, and video. Plus handles up to 3 hours of audio and 1 hour of video per session, with audio input in 90+ languages and speech output in 30+ languages across 55 voice timbres.
-

Opus 4.6 Sets Remote Labor Index Record, Questions AGI Claims
By
–
Holy smokes! Opus 4.6 set a new record on the Remote Labor Index! At 4.17%. Anyone who claims that we are close to AGI is either lying or lost.
-
LFM2.5-350M: Tiny Agentic AI Model Running in Browser
By
–
This is an always-on model living in your browser.
— Maxime Labonne (@maximelabonne) 31 mars 2026
Sub-500 MB QA, data extraction, tool use 🪄 https://t.co/6ttaJBTxAcThis is an always-on model living in your browser. Sub-500 MB QA, data extraction, tool use 🪄 Xenova (@xenovacom) NEW: LiquidAI just released LFM2.5-350M, a tiny model that brings agentic AI and tool-calling capabilities to resource-constrained environments. 🤯 It can even run locally in your browser via WebGPU, serving as a powerful companion while you browse the web. Try the demo! 👇 — https://nitter.net/xenovacom/status/2039043406823833964#m
→ View original post on X — @maximelabonne, 2026-03-31 21:46 UTC
-

H Company Releases Holo3, Outperforming GPT-5.4
By
–

H Company released Holo3, a new series of SOTA "Computer Use" models that outperform GPT-5.4 and Opus 4.6 on OSWorld-Verified and other benchmarks.
-
Reinforcement Learning Limitations on Fine-tuned Model Prompts
By
–
No, RL doesn't fix it. It merely makes e smaller for prompts present in the fine-tuning set.
-
LLMs and Code Generation Systems: Clarifying Autoregressive Architecture
By
–
1. I never said LLMs were not useful 2. Code generation systems are not strictly auto-regressive LLMs. They produce multiple outputs and pick the best ones. 3. Your argument is as if I said "perpetual motion is impossible" and you responded "meanwhile, it's been 300km since I
-
Autoregressive Models Error Propagation in Discrete Sequences
By
–
That's a ridiculous argument. – all auto-regressive models diverge, whether they are generative (in input space) or not. – for discrete symbol sequences, the probability of correctness decreases exponentially with the sequence length, assuming independence of errors. – THAT
-
Clarifying the utility debate around large language models
By
–
I never said LLMs were not useful. We're discussing a different question here.
