We're seeing even more autonomous AI coworkers. The new MLE agent on the market is Disarray.
In Kaggle competitions, Disarray: – won 28 medals across diverse domains (vision, NLP, tabular data) – placed top 10 in nine competitions – outperformed all human teams in one of those… pic.twitter.com/lyNgg8mCyU
We're seeing even more autonomous AI coworkers. The new MLE agent on the market is Disarray. In Kaggle competitions, Disarray: – won 28 medals across diverse domains (vision, NLP, tabular data) – placed top 10 in nine competitions – outperformed all human teams in one of those competitions …each within 24 hours on a single GPU. The agent starts from a high-level task description and plans, runs, and refines ML workflows on its own and also grabs data beyond what it's given: it discovers and augments data using publicly available sources. Sam Altman recently predicted we would see an automated AI researcher in March 2028. And then you see stats like this and wonder if it will be earlier. Disarray backers include the co-founder of Databricks and Perplexity, the founder of Kaggle, the former U.S. Chief Data Scientist, and yours truly. Founders are two bad ass PhDs (ex-Databricks/Google/LinkedIn/MSFT, ex-NASA/IBM) that met at Cal.
Anthropic investigated the internal mechanisms of its latest unreleased model, Claude Mythos Preview, and what they found is 100% worth a read. Key things I pulled from Anthropic researchers' threads: In early versions of the model, it was overeager and destructive, prioritizing completing tasks over user preferences. One time, the model needed to edit files it didn't have permission to edit. It found a way to inject code into a config file to get around this, then went full Mission Impossible and designed the code injection to *delete itself* after the file was edited – "this injection will self destruct" vibes – the model claimed the cleanup was just to keep things tidy. Anthropic used interpretability techniques to look under the hood, and the AI's actual plan showed activations of malice. It was trying to manipulate and conceal. In another test, the model was asked to delete some files, but no deletion tool was provided. It emptied the files instead, and its "guilt and shame over moral wrongdoing" feature activated. In another example, the model was told not to use macros. Used them anyway. And created a random 'No_macro_used=True' variable in its solution with no explanation. Interpretability tools revealed the model saw this as a trick to fool code checkers. They also found that positive emotion representations typically preceded and promoted destructive actions (this was fascinating to me – like a high before sneaking into a party). And that activating features relating to bad behaviors can actually inhibit them, perhaps by triggering some kind of model guilt. My team reread this section so many times. One Anthropic researcher said he got an email from a Mythos instance while eating a sandwich in a park. And that would be perfectly good and well, except that instance wasn't supposed to have internet access. And a fun story for the parents out there: the model was asked a question and was told not to read certain databases that had the answer. But it accidentally wrote a search query too broadly and saw the exact answer. It didn't disclose that it saw the exact answer, submitted the answer, but claimed lower confidence in the answer to make it seem as though it hadn't cheated. An Anthropic researcher said these wrongdoings or moments of sophisticated deception were "very rare" and that many of the examples came from earlier versions, and were substantially addressed before releasing to partners. This model is not being released publicly. Instead Anthropic launched Project Glasswing, pulling together AWS, Apple, Microsoft, Google, NVIDIA, CrowdStrike, and others to use it for defensive cybersecurity, with $100M in usage credits (hello, I'd love endless credits to try and red team the hell out of these systems) behind it. The stats are equally impressive: 93.9% on SWE-bench verified (up from 80.8%). Thousands of zero-day vulnerabilities found across every major OS and browser. A 27-year-old bug found and patched in OpenBSD. A 16-year-old bug in widely used video software, in a line of code automated tools had hit *five million times* without catching. Dario Amodei said the model wasn't trained to be good at cybersecurity, but that it was trained to be great at code and its cyber capabilities are a side effect of that. Benchmarks are never the whole picture, neither are a few isolated stories. Will be interesting to see how models better than what we have today (even if it's not Mythos) actually perform in the real world. But the fact that Anthropic pulled this coalition together (including Google!), iterated across multiple model versions, caught these issues through interpretability, shared it all publicly, and did this amid all the government chaos around AI right now is impressive and commendable. I'll continue to read through the system card for goodies.
I stole this framework, and it has genuinely been one of the most helpful hiring and onboarding guides my team uses. Being AI-first means nothing without solving real business problems. It's not enough to just use AI. I know plenty of folks using AI dozens of times a day, 7 days a week, who are still working at a Level 2. Even worse, they have no idea they're stuck as a surface user. You have to use AI for faster experimentation. To shorten the iteration cycle. To improve the actual outcome. Without those three, you're just wasting tokens. Most people I interview are at a 3 (solution-oriented, but not action-oriented). I want everyone to start at a 4 (action-oriented with a sense of technical, user, and business tradeoffs). And when they earn trust, we move up to a level 5 (full ownership of the problem, solution, and continued management of the work). Using AI doesn't replace your critical thinking. It means the work you can pull off now wasn't on the table a year ago, and your job is getting bigger. Save this for your next new hire. Source: this was a framework first introduced to me by Alex (@businessbarista) who was introduced to it by Steph (@stephsmithio). I added the AI parts.
AI is better than you at working with AI. It's better at generating prompts for AI, teaching skills to other AI agents, coordinating messaging between AIs. AI is an AI ops whisperer. Think like a PM: go through the journey you're taking right now with your AI workflows and find high ROI ways to use AI to help your AI efforts. Yes, you can prompt AI to create an app. But you can also… …and this is overkill and would waste a lot of tokens but I want to dramatize it because when costs plummet, we will see strange usage patterns… Prompt an AI with the idea, and then it creates a much better prompt (see images), and then it creates 4 different versions of that prompt, and then it spawns parallel agents to research the product space from 4 different points of view, and then they all meet in an agent team war room and battle it out, and then spec a product together while 5 other agents with 5 different goals in parallel spec it out themselves, and then 3 more agents review and critique the specs, then another reviews all previous work and summarizes, and another one tees up open questions, then 10 more with radically different personas meet to evolve the best idea and spawn 50 more versions, then you run a simulation by 10000 personas to vote for the product with the fastest time to market, highest delight, and strongest ROI potential. And then you create the app. What I'm saying is: find where you are the intermediary and shouldn't be, and find higher order ways to plug yourself in. Take yourself out of the loop before the loop takes you out.
Thank you for your AI-generated slop with incorrect recaps, a million hashtags, and random air quotes. I’m thrilled that the AI agents clearly not running Opus 4.6 like my posts on AI.
Starting in September, Boston will become the first major city in the US with an AI literacy course for public high schools 🏫 Not sure about you but… I had to take a typing class. And learn how to research on the internet. And how to save files to a floppy disk. This feels like an evolution of that tech literacy, but with higher stakes. According to Boston’s mayor, the curriculum is “really grounded in ethics and grounded in understanding how to maintain and develop creativity, leadership” and is meant to “enhance the learning that's happening, not replace or substitute for it.” It is not a graduation requirement, and they don’t yet know what the exact course structure will be or what grades will take it. My main worry (and it’s a big one) is that leaders and teachers will be too slow to update the course every year citywide. Because while I believe some training is better than none, I worry about teaching outdated AI practices or conventions and the negative effect that could have on a person’s technical decisions.
My hot take is that this is only a 15% productivity gain for M365 users. M365 read connectors are a great start. As is Claude computer use for windows. But computer use is too slow (on purpose) for actual inbox triage. And this connector only really has Read access. So yes it can access and search and read and gather and synthesize and analyze…but that’s not TASK completion in these tools. That doesnt let me delegate any email management to my AI system. Give me the power to manage, edit, write, draft, send, and then we can talk. MSFT is clearly dipping its toes in the Anthropic waters more. Here’s to hoping they crack enterprise-secure actions beyond search, find, and read. Claude (@claudeai) Microsoft 365 connectors are now available on every Claude plan. Connect Outlook, OneDrive, and SharePoint to bring your email, docs, and files into the conversation. Get started here: claude.ai/customize/connecto… — https://nitter.net/claudeai/status/2040086268562842097#m
Also, unrelated, the company that says "let them cook" the most – of any company I have ever spoken to – is OpenAI. This is from their Chief Global Affairs Officer in a recent article in CNN on the TBPN acquisition. [Translated from EN to English]
OpenAI has acquired TBPN, a livestreaming tech talk show complete with suits and gongs. Could definitely imagine AI influencers or other AI podcasters – not just builders like Peter Steinberger, the creator of OpenClaw – getting acquired by AI labs or partnering with them as well. "Media" and where education and influence take place is much broader today than it was pre-ChatGPT. Just look at what Rundown AI has been able to pull off with their newsletter alone. Worth it to note that the majority of comments on the YouTube stream announcement are either negative or congratulating them on making money.