Imagine seeing how capable Qwen 3.6 27B is when you give it web access and a proper harness and not yet getting that AGI will run locally Ngmi
AI
-
Evolution of AI Models and Benchmark Shifts
By
–
models will undoubtedly get to that point, and the METR benchmarks will undoubtedly shift to a frame above their current one—of which there are many
-
Codex Automates Expense Reimbursements
By
–
Codex quite literally filed my reimbursements, downloaded invoices since the start of the month, updated the expenses spreadsheet and filled the actual form all by itself Used Drive & Sheets plugin for state tracking
Gmail plugin for tracking invoices
Chrome extension for actual -
Early Look at Imagine Agent Mode for Image and Video Generation on Grok App
By
–
Early look at Imagine Agent Mode on Grok app for iOS!
— 🚨 AI News | TestingCatalog (@testingcatalog) 9 mai 2026
Users will be able to use Imagine Agent via a mobile optimised native UI to generate images and videos that require more complex workflows.
SpaceXAI is getting quite ahead of everyone else on this front!
We just need… pic.twitter.com/5QxeCclHEoEarly look at Imagine Agent Mode on Grok app for iOS! Users will be able to use Imagine Agent via a mobile optimised native UI to generate images and videos that require more complex workflows. SpaceXAI is getting quite ahead of everyone else on this front! We just need
-
Critiquing the practice of testing AI tools for failure
By
–
“We got a tool to perform poorly” is the lowest form of science and journalism imo and is only relevant when the tool is, in fact, extremely useful
-

Continuous Latent Diffusion Language Model Advances
By
–
“Continuous Latent Diffusion Language Model” Most diffusion language models still use diffusion to recover token-like states, just in a different generation order. However, this paper uses diffusion in a different way. It learns a continuous latent prior for global semantics
-

Sparser, Faster, Lighter Transformer Language Models
By
–
“Sparser, Faster, Lighter Transformer Language Models” LLMs are naturally sparse in their feedforward layers, but unstructured sparsity usually doesn’t get you real speed on GPUs, because the hardware stack is built for dense compute. The key idea of the paper is to redesign
-
AI Capable of Independent Organizational Work is Coming Soon
By
–
> Level 5 > Organizations
> AI that is capable of doing all of the work of an organization independently Soon -

How Human Prompting Influences AI Model Benchmarks
By
–
mythos obviously looks incredibly capable and im psyched to use it also if you're panicking about it: benchmarks don't measure model capability alone they measure model capability after a human has done the work of finding a prompt that lets the model’s capability appear that
-
Using LangSmith for Organizational AI Agent Collaboration
By
–
one way to view langsmith is as a platform for the whole org to collaborate on building agents helps speed up that feedback loop between different personas