RL keeps evolving! Now you can teach LLMs to reason better by rewarding risk-taking. Risk-based Policy Optimization (RiskPO) is a new reinforcement learning framework for post-training LLMs. Instead of averaging rewards like GRPO, RiskPO uses a Mixed Value-at-Risk objective
LLMS
-
Windsurf launches SWE-grep for faster context retrieval
By
–
Windsurf announced SWE-grep and SWE-grep-mini models for faster context retrieval. These models are now available on Windsurf!
— 🚨 AI News | TestingCatalog (@testingcatalog) 16 octobre 2025
TPS 🤯 https://t.co/E3QgCw8ndK pic.twitter.com/VMWuFaUDRrWindsurf announced SWE-grep and SWE-grep-mini models for faster context retrieval. These models are now available on Windsurf! TPS
-
GPT-6 Likely Arriving Before Year End, CNBC Reports
By
–
BREAKING 🚨: GPT-6 is likely coming before the end of the year! This means, less than 2.5 months!
— 🚨 AI News | TestingCatalog (@testingcatalog) 16 octobre 2025
From the "Rise of agentic commerce could dramatically change the tech landscape: Evercore ISI's Mark Mahaney" video by CNBC
"Wen GPT-6?" time has come 👀 https://t.co/IxYIYTtwm8 pic.twitter.com/X16dyQIjHLBREAKING : GPT-6 is likely coming before the end of the year! This means, less than 2.5 months! From the "Rise of agentic commerce could dramatically change the tech landscape: Evercore ISI's Mark Mahaney" video by CNBC "Wen GPT-6?" time has come
-
GPT-5 Outperforms Sonnet for Idea Brainstorming Tasks
By
–
ChatGPT is the way to go! Specially for bouncing ideas – gpt-5 is way more steerable than sonnet
-
1B Parameter Model: 128K Context, Int4 Quantization, Llama 4 Distilled
By
–
chat is this real??? 128K context, int4 quantisation, 1B params, distilled from Llama 4
-

LLM-as-Judge Alignment Improves to 94% With Structured Evaluation
By
–
After refining our rubric design process, alignment between human reviewers and LLM-as-a-Judge improved from 37 % → 94 % — proof that structured evaluation enhances consistency.
-
Rubric-Driven Evaluation Improves LLM Alignment and Data Quality
By
–
Rubric-driven evaluation doesn’t just improve human and LLMaJ alignment. It accelerates delivery, reduces rework, and improves the thing we care about most: data quality.
-
Trusted Scale: Snorkel’s Framework for Scaling Trust in AI
By
–
This is Trusted Scale — Snorkel’s framework for designing, validating, and scaling trust in AI data pipelines. Read the full post by @pham_derek →
https://
snorkel.ai/blog/scaling-t
rust-rubrics-in-snorkels-quality-process/
… #AI #LLM #DataQuality #SnorkelAI #MachineLearning -

GPT-5 o3.1 iteration represents next leap in AI reasoning capabilities
By
–
Lead OpenAI researcher: "GPT-5, in some way, can be considered o3.1 – iteration of the same concept"
— Peter Gostev (@petergostev) 16 octobre 2025
"What I'm after right now is something next, what would be a significant jump to how we interact with models, that are more capable, think for even longer and interact with even… https://t.co/bZOtxQYXHYLead OpenAI researcher: "GPT-5, in some way, can be considered o3.1 – iteration of the same concept" "What I'm after right now is something next, what would be a significant jump to how we interact with models, that are more capable, think for even longer and interact with even
-
Poe Leaderboard Now Available for Desktop and Web
By
–
To view the current rankings, go to https://
poe.com/leaderboard (currently desktop and web only). (2/2)
