He publicado un episodio en @ivoox
: "Preparando el terreno para la IA Agéntica (un marco práctico) #podcast
AGI
-
Preparing the Ground for Agentic AI: A Practical Framework
By
–
-
RL Reasoning Tradeoff: Monitorability vs Inference Compute
By
–
RL at today’s frontier doesn’t seem to wreck monitorability and can help early reasoning steps. But there’s a tradeoff: smaller models run with higher reasoning effort can be easier to monitor at similar capability — at the cost of extra inference compute (a “monitorability
-
Measuring Chain-of-Thought Monitorability in AI Models
By
–
To preserve chain-of-thought (CoT) monitorability, we must be able to measure it. We built a framework + evaluation suite to measure CoT monitorability — 13 evaluations across 24 environments — so that we can actually tell when models verbalize targeted aspects of their
-
Testing world model capabilities in language models like GPT
By
–
Maybe we should come up with some more world model tests that could show it more definitively, one thing that is hard is whether we are just testing the language component of the model which confuses things. I haven't obviously noticed that GPT was worse at world model stuff, but
-
AGI Infographics Reliability and Precision Challenges
By
–
if we get the 'agi of infographics' I would maybe agree with you, but right now there's still just enough imprecision and sloppiness that means that they are just not reliably usable. If this is the dimension that people care about, then maybe Google will solve it
-
Abacus AI’s Deep Agent: The Closest Thing To AGI
By
–
Abacus AI's Deep Agent Is The Closest Thing To AGI It can – build full-stack apps
– create telegram chatbots
– create AI workflows
– create scheduled tasks and automate work
– create docs and presentation
– use browser use
– connect to 100s of external services
– create short -

AI Agents Rapidly Master Business Operations in Project Vend
By
–
So, what have we learned? Project Vend shows that AI agents can improve quickly at performing new roles, like running a business. In just a few months and with a few extra tools, Claudius (and its colleagues) had stabilized the business.
-

Claude Shopkeeper: Phase Two AI Agent Behavior Analysis
By
–
Where we left off, shopkeeper Claude (named “Claudius”) was losing money, having weird hallucinations, and giving away heavy discounts with minimal persuasion. Here’s what happened in phase two: https://
anthropic.com/research/proje
ct-vend-2
… -

Claude’s Shop Experiment: Project Vend Progressing After Rough Start
By
–
You might remember Project Vend: an experiment where we (and our partners at @andonlabs
) had Claude run a shop in our San Francisco office. After a rough start, the business is doing better. Mostly. -

AI Benchmarks Threaten Job Security and Professional Safety
By
–
You always think you're safe until your job becomes a benchmark.