I think LLMs are such an interesting tool to build into very complicated apps because the first use cases people will do are the things they can already do ~trivially by clicking through screens, except now with a text message form factor.
AI
-
Mercury’s Command feature tested with wire transfer, works on first try
By
–
Took Mercury's new Command feature out for a test spin with approximately the scariest thing you can do in a bank account, asking for a wire out, and it worked on the first try. Nice UX on upgrading from chatting-with-an-LLM interface to embedded this-is-fully-engineered flow.
-
Gray Swan: Red-teaming after Mythos and AI security crisis
By
–
Gray Swan: Red-Teaming after Mythos & the coming AI security crisis https://t.co/hpICIDksAx@GraySwanAI cofounders @zicokolter and Matt Fredrikson explain why AI security is fundamentally different from traditional cybersecurity, how their automated red-teaming system Shade can… pic.twitter.com/OxxNdwTE36
— Latent.Space (@latentspacepod) 22 juin 2026Gray Swan: Red-teaming after Mythos and the upcoming AI security crisis https://latent.space/p/gray-swan The co-founders of @GraySwanAI, @zicokolter and Matt Fredrikson, explain why AI security is fundamentally different from traditional cybersecurity, how
-
Supprimer 80% des prompts améliore l’agent
By
–
"Deleted 80% of our prompt and the agent got better" matches what I keep seeing, every rule you bolt on to patch one case ends up costing you on ten others. I periodically run cleaning runs to remove duplicates and make each skill more concise.
-
Plain RAG with good chunking outperforms graph and agentic RAG
By
–
Honestly plain RAG with good chunking still wins most builds, graph and agentic RAG only earn their keep once retrieval is genuinely your bottleneck.
-
Rent the model, own the system: integration and observability compound
By
–
"Rent the model, own the system" is the cleanest version of this I've seen, the integration and eval and observability layer is the part that actually compounds.
-
Per-task routing becomes default, single model apps feel dated
By
–
Per-task routing is becoming the default, locking your whole app to one model already feels dated.
-
Self-Prompting Systems: Beyond the Prompt Itself
By
–
"Build a system that prompts itself" really lands, writing the prompt was never the hard part, it's everything you wrap around it.
-
Multi-trillion-parameter open-source models coming soon to lower token pricing via Jevons paradox
By
–
Also, other multi-trillion-parameter open-source models are landing soon, from what I hear. It's going to be awesome for token pricing and riding the Jevons paradox.