two fun surprises from using activegraph:
– the coding agent i was using would query the trace db to debug instead of looking at the logs like they normally would (i didn't ask it to)
– when long eval runs broke (laptop, api, etc.), it was always able to pick up from right before
AGENTS
-
Two fun surprises from Activegraph: agent uses trace DB, evals resume
By
–
-

Gated approach to agent self-modification via forking and testing
By
–
less novel, but still very interesting impo is the gated approach to self-modification the agent basically forks itself, propose a patch, run through multiple tests (static/sandbox/diff), and something called a binding held out gate before modificaiton lands
-

Showcasing controlled self improvement with regime-to-seam approach
By
–
i showcase "controlled" self improvement with a novel regime-to-seam approach where failures are categorized and allowed to fix targeted areas of the agent while interesting, it's more to showcase the type of self-modification that's easy to set up with activegraph
-

ActiveGraph: Auditable Gated Improvement Loop Demonstrated on LongMemEval
By
–
in arxiv paper #2, i tackle the last topic from paper #1: @activegraphai as an architectural affordance for self-improving agents "Regimes: An Auditable, Held-Out Gated Improvement Loop Demonstrated on LongMemEval with ActiveGraph" i demonstrate this with a reproducible gated
-
Open Architecture Agent Blueprint of Nebius AI with LangSmith
By
–
The open reference architecture of the @nebiusai Agent Blueprint connects proven components at every layer of the agent stack. We are thrilled to have Deep Agents and LangSmith as integral parts. Full announcement.
-

AI agents improve their own control harnesses
By
–
—
"Auto-Harness: Harnesses That Improve Themselves" What if an AI agent improved the harness that controls how it acts? Thus, instead of humans adjusting prompts, tools, retry rules, and verification for each model, this article explores
— -

Robinhood agentic trading account opened, plans automation with Hermes Agent
By
–
Woah…Just opened an Agentic Trading Account with Robinhood. Codex and Claude Code (Fable 5) can now run wild with it. After sometime I plan to automate this end-to-end with Hermes Agent.
-

Evaluating AI Agents Bootcamp June 27 with discount code KIRK50
By
–
Evaluating AI Agents — Bootcamp on June 27 hosted by @PacktPublishing @PacktDataML — register with my discount code 'KIRK50' at this link to save: https://
eventbrite.co.uk/e/evaluating-a
i-agents-bootcamp-tickets-1990306501323?aff=kirk
… Participants will be guided through the fundamentals of Evaluating AI Agents, covering: Why modern agents -

Stop prompting coding agents, design loops instead – loop engineering
By
–
Stop prompting your coding agents.
The people building Claude Code and Codex don't prompt anymore — they design loops. New video: what loop engineering actually is, the 5 building blocks, and the 2 traps that burn your token budget. https://
youtu.be/NjXIIH9vcv0
