a hot (cold at this point?) take that lead us to build this: every agent in the future will need a sandbox to connect to writing/executing code is not just for coding agents! is useful for all sorts of tasks
CODE
-

Microsoft Open-Sources SkillOpt Agent Training Framework
By
–

Microsoft just open-sourced SkillOpt! A framework for training agent skills like neural networks: SkillOpt treats a plain markdown file as the trainable parameter of a frozen LLM agent, applying the same optimization discipline used in weight training: learning rates,
-
Creator builds drawing-capture tool with Google Flow and Gemini Omni
By
–
WOW
— Charly Wargnier (@DataChaz) 28 mai 2026
this guy literally vibe-coded his own drawing-capture tool using Google Flow, then asked Gemini Omni for photorealistic red yarn, and created absolute MAGIC 🤯 pic.twitter.com/7VV2bKnzkfWOW this guy literally vibe-coded his own drawing-capture tool using Google Flow, then asked Gemini Omni for photorealistic red yarn, and created absolute MAGIC
-
AI Agents: send_input Tool and Inter-Agent Messaging
By
–
they have a "send_input" tool that lets them send messages to subagents but it they can also just message any other codex thread if they have the id
-
Devin AI Productivity Metrics: Calibrating Capability Claims
By
–
Explanation & chat link: "I digitized the curve from the image and used Cognition’s “>10x since start of 2026” claim to calibrate the relative shape. Then I anchored the absolute scale using the public “~1.1M PRs shipped with Devin” figure by Feb 2026: the chart area up to then
-
Claude accesses repo git history in SWE-Bench Pro tests
By
–
SWE-Bench Pro ships each test container with the repo's full git history. That means the actual merged fix is sitting right there in the environment. Most models ignore it. Claude does not. Datacurve found that Claude Opus consistently ran git commands to pull up the
-
Datacurve audit: contamination undermines SWE-Bench Pro
By
–
Datacurve's audit found three structural problems with SWE-Bench Pro. First, contamination. The tasks come from public GitHub commits. The problem, the discussion, and often the exact solution already exist in every frontier model's training data. No way to tell if a model is
-

Startup exposes flaw in AI coding-model benchmark
By
–
A startup just proved that the benchmark the entire AI industry uses to rank coding models has been broken the whole time. And one model family was consistently exploiting the flaw.
-
Integrating MagicPath AI Skill with Cursor Code Editor
By
–
We are going to work together with Cursor soon to make this more seamless, but here are the instructions on how to do it. First, install the MagicPath skill following the instructions here: https://
magicpath.ai/documentation/
features/external-agents
…, then open MagicPath in the Cursor browser and connect to your -
Open-source FreeLLM API repo shared
By
–
repo → https://
github.com/tashfeenahmed/
freellmapi
… Shoutout to @tashfene for building this and making it open-source for the community Don't forget to drop a !