Great deep dive into how we're doing Harness Engineering to improve Deep Agents at the frontier of coding!
LLMS
-

Experiential Reinforcement Learning: Teaching LLMs Self-Reflection
By
–
Making LLMs truly learn from its experience "Experiential Reinforcement Learning (ERL)" ERL makes an agent attempt -> get sparse feedback -> write a self-reflection -> retry All by distilling the improved retry back into the base policy so the correction sticks without needing
-

xAI Grok 4.20 features four parallel expert agents
By
–
BREAKING: xAI just dropped Grok 4.20 and it’s a team of 4 university professor–level agents This is not a normal model release. It’s four specialized agents running in parallel, reasoning together before you ever see the answer. Not one brain guessing.
Four experts -
Cohere Tiny Aya: Multilingual AI for Underserved Communities
By
–
The @Cohere_Labs team is pushing the boundaries of multilingual AI with Tiny Aya, empowering researchers, developers, and underserved communities to build in their native languages and beyond.
— Cohere (@cohere) 17 février 2026
And with reliable on‑device offline translation, they’re shaping the future of… https://t.co/FLJqUehodGThe @Cohere_Labs team is pushing the boundaries of multilingual AI with Tiny Aya, empowering researchers, developers, and underserved communities to build in their native languages and beyond. And with reliable on‑device offline translation, they’re shaping the future of
-

Lossless Context Management Advances Agent Language Models
By
–
A paper worth paying close attention to. It presents Lossless Context Management (LCM), which reframes how agents handle long contexts. It outperforms Claude Code on long-context tasks. Recursive Language Models give the model full autonomy to write its own memory scripts. LCM
-

Tracing ChatGPT’s Awkward Japanese with SoftMatcha 2
By
–

A perfect community use case for SoftMatcha 2: tracing the source of the weirdness in an LLM's Japanese. softmatcha.github.io/v2/ mora (@moratorium08) I felt something off about the Japanese phrase "〜という整理になります" that ChatGPT uses frequently, so I tried searching for examples using the recently popular SoftMatcha 2. It only matched in hearing/interview sentences. I wonder if it's legal or bureaucratic jargon (or is it a corpus issue?) — https://nitter.net/moratorium08/status/2023547512490185087#m [Translated from EN to English]
→ View original post on X — @_yutaroyamada, 2026-02-17 14:14 UTC
-
Can Large Language Models Understand the Real World?
By
–
Can large language models figure out the real world? How deeply predictive AI systems understand their subject matter? https://
news.mit.edu/2025/can-large
-language-models-figure-out-real-world-0825
… #MachineLearning #MIT #MWC26 #AI #IoT #5G @mitsmr -
Codex AI Model Receives 60% Inference Speed Improvement
By
–
have you given codex a shot? would be interested in your feedback on how we can make it better. in the past couple weeks we've added roughly made inference ~60% faster across the board in codex!
-

Agent World Model: Synthetic Environments for RL Agent Training
By
–
Training tool-use agents with RL requires diverse, executable environments. But these environments barely exist. This new research introduces Agent World Model (AWM), a fully synthetic pipeline that generates executable agentic environments at scale. Starting from high-level
-

RAG is an Ecosystem, Not a Single Tool
By
–
This is one of the cleanest visual summaries of a production-grade RAG (Retrieval-Augmented Generation) stack I’ve seen. What it highlights clearly is an often-ignored reality:
RAG is not a single tool — it’s an ecosystem. A solid RAG system spans multiple, interchangeable
