Paper does test multi-turn (24 tool calls in OfficeQA, 30 turns in SpreadsheetBench, 50 steps in ALFWorld), but mid-session auto-compaction isn't part of the eval. Skill being under 2K tokens probably helps it survive compaction, but not validated against that failure mode.
AI
-
Optimization Stack in AI: From Weights to Skill Files
By
–
YES, Optimization keeps moving up the stack. Weights, then prompts/harness, now skill files.
-

Hundred-Page Language Models Book with PyTorch by Burkov
By
–
The Hundred-Page Language Models Book — Hands-on with PyTorch: http://
amzn.to/4sJl7YC by @burkov -
Prompt Engineering with Omni for Video Generation
By
–
One year later with Omni and this test can pass. I saw it getting pretty close, so I tweaked the prompt:
— fofr (@fofrAI) 26 mai 2026
> A video of a man counting to 10 on his fingers, show the number in the corner. A new number every 1s, no dialogue other than the numbers he says. He uses two hands for… https://t.co/TPd8HEl97W pic.twitter.com/09M1eYw0qVOne year later with Omni and this test can pass. I saw it getting pretty close, so I tweaked the prompt: > A video of a man counting to 10 on his fingers, show the number in the corner. A new number every 1s, no dialogue other than the numbers he says. He uses two hands for
-
AI Internal States Mirror Human Neuroscience Findings
By
–
> … [W]e keep finding things that are mysterious, even unsettling. We find structures that mirror results from human neuroscience. We find evidence of introspection. We find internal states that functionally mirror joy, satisfaction, fear, grief, and unease. I don’t know what
-
AI turns past prompts into reusable skills and subagents
By
–
It will look through your past sessions and turn repeated prompts into reusable skills + subagents
-
Troubleshooting AI behavior by removing persistent context
By
–
I do use structured prompts, plan/goal mode, Notion context, and a few local reference docs. Memory is involved. Could be instruction/memory interference. I’ll try a clean run without persistent context and see if it changes the behavior. thanks!
-
GPT instruction following worsened, hallucinations and errors compared to 5.5
By
–
both id say. Instruction following has worsened. GPT repeatedly hallucinates during the task, then admits mistakes, tries to correct them, and makes further serious errors. I haven't encountered such serious errors with version 5.5.
-
Turn Claude into Specialist Marketing & Business Agents
By
–
If you find this useful, also check this out. Turn Claude into 20+ different specialists for marketing & business. Install real expertise, not just prompts. Get my Claude skills bundle
-
Automated 30-day review of recent AI work using memory and git
By
–
The text version of the prompt: "Look back over my recent work from the last 30 days using all available context. Use available evidence in this order:
– Session Memory summaries and MEMORY. md entries
– Git log and recent commit history across branches
– CLAUDE. md and