That’s the one I am still finalizing, hopefully will have emails sent tomorrow for attaching proofs and a few days after will have a randomly picked name announced
CODE
-
Local Models Excel at Code and Tools Tasks
By
–
Some of the latest local models appear to be much better with code and tools than just a few months ago, so I buy that you need to carefully pick the right model+harness combo in order to see them at their best
-

Must-Read AI Research of the Week: LLM Agents and Optimization
By
–
Must-read AI research of the week: ▪️ Learning to Commit: Generating Organic Pull Requests via Online Repository Memory ▪️ Effective Strategies for Asynchronous Software Engineering Agents ▪️ Composer 2 ▪️ From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents ▪️ Scalable Prompt Routing via Fine-Grained Latent Task Discovery ▪️ MSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT ▪️ On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation ▪️ Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs ▪️ Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? ▪️ RL for Distributional Reasoning in LMs ▪️ Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought ▪️ EVA: Efficient Reinforcement Learning for End-to-End Agent Find the full list and the main AI news here: turingpost.com/p/fod146
→ View original post on X — @debashis_dutta, 2026-03-30 23:41 UTC
-
Development of AI Tooling and MCP Server Integrations
By
–
good question, I started by contributing tooling around http://
developers.openai.com with llms.txt support and then added the Docs MCP server Supported myriad of model launches from OpenResponses, GPT 5.3 Codex, Spark, GPT 5.4 Along with Codex App and features like Skills, Sub -

AI Coding Capabilities Improving: Developer Skepticism Addressed
By
–
This is an insane Anthropic tweet. And it’s a *buried reply* to one of their other tweets. I am reminded of a talk I gave ~2-3 months ago where a senior developer at a Fortune 500 company asked me “why would I use AI to code if I can just code myself.” I answered. He said, “But sometimes it messes up.” I told him this was coming. Even if it’s not perfect today (it makes weird product features decisions sometimes, not gonna lie), the scaling laws seem to be holding up this year and the next iteration will be even more capable. I wish I could send him this tweet.
→ View original post on X — @alliekmiller, 2026-03-30 21:26 UTC
-

Natural Language Agent Harnesses: From Code to AI-Defined Control Logic
By
–
We’re trying to build intelligent systems… using control frameworks designed by humans. That’s the core limitation of today’s agent harnesses. A new paper from Tsinghua University and Shenzhen proposes something radically different: 👉 What if the harness itself is not code—but natural language? Instead of hardcoding orchestration logic, they introduce Natural-Language Agent Harnesses (NLAH): – The control logic is written as an editable natural language SOP – The LLM interprets and executes that SOP dynamically – A shared runtime enforces structure via contracts, artifacts, and adapters Even more interesting: ➡️ The SOP itself can be generated and adapted by AI depending on the task So instead of: > Humans define → Agents execute We get: > AI defines → AI executes → AI evolves 🧠 Technical takeaway This shifts agent design from: – Static orchestration graphs – Hardcoded tool pipelines – Rigid planner-executor loops To: – Executable natural language control logic – Runtime-interpreted orchestration – Portable, composable harness artifacts The harness is no longer buried in code—it becomes a first-class abstraction. 🏗️ Architecture implications – Decouple control logic from implementation – Treat orchestration as data, not code – Use LLMs as meta-execution engines – Design systems that scale with tokens, not constraints 💡 Bigger question If agents can define and execute their own control logic… What else in AI system design should stop being code—and start being language? 📄 Paper: arxiv.org/abs/2603.25723 🔗 Follow my communities and personal initiatives: • Amazing AI, Data, Quantum Computing & Emerging Technologies — drdebashisdutta.com/ • Research & Innovation – Quantum, AI & Advanced Systems — researchedge.org
→ View original post on X — @debashis_dutta, 2026-03-30 21:25 UTC
-
Abacus CoWork Launches Multi-Model AI for Laptops
By
–
🚨 BREAKING NEWS – Abacus CoWork brings Claude, GPT 5.4 And Gemini To Your Laptop!
— Bindu Reddy (@bindureddy) 30 mars 2026
Super excited to announce our MULTI-MODEL CoWork product!
– combine the coding power of Opus with the reasoning prowess of GPT 5.4
– optimized for efficiency using "low effort" mode
– computer… pic.twitter.com/ppUSqjIWbs🚨 BREAKING NEWS – Abacus CoWork brings Claude, GPT 5.4 And Gemini To Your Laptop! Super excited to announce our MULTI-MODEL CoWork product! – combine the coding power of Opus with the reasoning prowess of GPT 5.4 – optimized for efficiency using "low effort" mode – computer use to tes – packaged for FREE with ChatLLM and Abacus AI's Deep Agent Get complex tasks done right on your laptop
-

GRASP: New Gradient-Based World Model Planner Released
By
–
Code for our new world model planner is live! github.com/michael-psenka/gr… Includes our implementation on dino-wm, as well as implementations on jepa-wm and le-wm, and minimal pseudocode for anyone to re-implement themselves. Michael Psenka (@michaelpsenka) tl;dr New planner for world models! GRASP: gradient-based, stochastic, parallelized. Long range planning for world models has always been an issue. 0th order methods like CEM/MPPI dominate, but have degrading performance at longer contexts or higher-dimensional actions. We wanted to address this from the ground up. w/ Michael Rabbat, @ask1729 , @ylecun*, @_amirbar* (equally advised) — https://nitter.net/michaelpsenka/status/2019870377032503595#m
→ View original post on X — @berkeley_ai, 2026-03-30 20:44 UTC
-
The rise of vibe coding in AI-driven development
By
–
and all is vibe coded, it's incredible… and computers can't do the coffee for dev
-
SambaNova Paper on Long-Context Reasoning Limits Accepted at ICLR 2026
By
–
Big moment for the team at SambaNova 🦾 Our paper has been accepted at ICLR 2026 The Limits of Long‑Context Reasoning in Automated Bug Fixing – even with 64k‑token windows, GPT‑5‑nano solves 0 %, Qwen3‑Coder 7 %. Success still comes from short‑step decomposition, not raw context size. arxiv.org/abs/2602.16069?utm…
→ View original post on X — @sambanovaai, 2026-03-30 20:30 UTC