Too much of the solutions & language for agentic systems comes from a coding perspective (control planes, hooks, loops), but I really think that the field of management and organizations can tell us more about how to work with agents (boundary objects, spans of control, etc).
@emollick
-
AI Challenges in Multi-Agent Systems and Organizational Workflows
By
–
For individual AI use, the jagged frontier is increasingly well understood. In multi-agent workflows in organizations, AI is jagged in ways that have not been well identified yet. In fact, we don't even have a vocabulary around multi-agent systems & the ways the fail or succeed.
-
GPT-5.5 Launch and Claude Finance Briefing on May 5
By
–
May 5 is the GPT-5.5 launch celebration in San Francisco and the Claude Finance Briefing in New York. Real opposite valence events on opposite coasts.
-
AI Regulation Criteria Vague Without Better Non-Lab Benchmarks
By
–
That doesn't mean that there should not be regulation and vetting, but it does suggest that is hard to write a criteria right now that is not somewhat vague. More R&D into non-lab benchmarks is urgently needed. We have remarkably few good ones that are unsaturated and clear.
-
AI Regulation Hampered by Poor Benchmarks and Risk Metrics
By
–
A challenge with AI regulation and vetting is how bad our benchmarks of AI model performance and risks are. There is no benchmark for risks and red-teaming requires experiments from dedicated specialist organizations & is not easy to put metrics around. No clear objective numbers
-

GPT-4o and Llama 3.3 Pose Low Harm Risk Study Suggests
By
–

I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & more sycophantic) chatbots essentially did nothing for people who followed their advice, it means that there is less risk of harm as well.
-

Retracted AI Education Paper Prompts Discussion of Meta-Analyses
By
–
My surprise here seems warranted, this paper was retracted (There are other peer-reviewed meta-analyses of the impact of AI on education finding positive effects, like: https://
researchgate.net/publication/38
7110151_The_effects_of_GenAI_on_learning_performance_A_meta-analysis_study
… though the best evidence of AI helping is from RCTs of interventions with AI tutors) -

Claude’s Internal Thought Trace on a User Request
By
–
Claude's thought trace on one such request – super interesting.
-

Poems LLMs Seem to Favor About Their Own Existence
By
–



Poems that ChatGPT, Claude, and Gemini all seem to "like" when you ask for poetry related to being/making LLMs:
Rilke's "Archaic Torso of Apollo"
Stevens' "Idea of Order at Key West"
Borges's "The Golem" (or "The Other Tiger")
Pessoa's "Autopsychography" Pretty apt choices! -
GPT-5.5 Pushes Back on Goofy Cover Letter Demo Requests
By
–
Sometimes when I demo AI, I show it turning cover letters into goofy formats (poetry, etc) as an introduction to the idea of AI as translator between forms. For the first time, GPT-5.5 has been trying to get me to tone these requests down so I don’t ruin my chances at the job.