Google just built an AI that organizes itself. It’s called TUMIX, and it might be the most interesting paper Google has published this year. Instead of training a bigger model, the team built a system where multiple AIs work together at test time. Each agent uses different
LLMS
-
LLMs Memory Security: Urgent Cognitive Security Improvements Needed
By
–
from what i have seen on this site, y’all really need to be upping your cogsec these llms will rewrite your memories if you don’t
-

Insulting LLMs: Unexpected Prompt Optimization Technique
By
–
Ceux que j'ai formé en sont étonné, mais oui insulter un LLM donne parfois des meilleurs résultats. J'ai expliqué en commentaire pourquoi. …et c'est a la dernière fois qu'il vous donne le meilleur résultat. Vous voulez avoir un meilleur résultat en 1er ? Ne me croyez pas
-
Claude Code Interpreter Skills Public Folder Discovery
By
–
I just learned Claude's new code interpreter mode has a /mnt/skills/public/ folder full of prompt instructions and Python utilities for creating and manipulating pdf, docx, pptx, and xlsx files – and you can ask Claude for a copy and learn a TON about working with those formats
-

Open Source Model Inference Prices Drop Significantly Over Time
By
–
Something I haven't thought about too much, but inference prices for open source model shift quite a bit over time – e.g. I've taken 6 models for which I had snapshots of data in August and October 2025 and variance has been quite large: ~30% drops for OpenAI models and ~60%
-
Running GPT-OSS 20B on Snapdragon Phones with GPU Memory
By
–
TIL you can run GPT-OSS 20B on a phone! This is on Snapdragon phones with 16GB or more of GPU-accessible memory – I didn't realize they had the same unified CPU-GPU memory trick that Apple Silicon has (The largest iPhone 17 still maxes out at 12GB, so not enough RAM to run https://
x.com/nexa_ai/status
/nexa_ai/status/1975232300985291008
… -
Dishonesty in Benchmark Reporting: A Corporate Risk
By
–
It's a useful skill but it's still a red flag if you are dishonest about it. Imagine asking the person to report benchmark performance of the LLM that is being developed, would you trust the results? This can backfire on the whole company.
-

Seeking GLM Researcher Speaker for AIE CODE Conference
By
–
Do I know anyone affiliated with GLM? we are looking for one last speaker for AIE CODE and a GLM/QwenCoder/Kimi researcher is probably the right choice here
-
Last Chance: Context Engineering Webinar with LangChain and Manus
By
–
Last Chance to Join: Context Engineering with LangChain & @manusai This webinar brings together @RLanceMartin Martin (Founding Engineer, LangChain) and Yichao “Peak” Ji (Co-founder & Chief Scientist, Manus; MIT Innovator Under 35) to share practical best practices in
-

Smaller Models Gain Accuracy Through Multi-Agent Coordination Setup
By
–
In the switch to the multi-agent setup: Smaller models regained lost accuracy when configured as coordinated agents Long-context reasoning improved
