It would also be useful for funding R&D into benchmarking models, which is currently mostly done by the labs themselves right now.
RESEARCH
-
NIST Public AI Ability Tests as Independent Evaluator
By
–
In addition to the CAISI evaluation, it would be useful if NIST conducted public tests of AI abilities as an independent evaluator – though those obviously should not be pre-release tests & can be done when models are public. Independent testing is important & getting expensive.
-

New Microsoft Research Paper on Long-Horizon Agent Generalization
By
–
NEW paper from Microsoft Research. Nice study on long-horizon agent generalization. (bookmark it) The team runs a study where the only variable is task horizon length. They use the same decision rules, reasoning structure but different sequence length to the goal. The main
-
SubQ vs Opus 4.6: massive context, cheaper, faster, threatens Claude
By
–
Si los números de SubQ son reales, Claude tiene un problema serio.
— Nico (@nicos_ai) 5 mai 2026
SubQ vs Opus 4.6:
→ 12M tokens de contexto (Claude se rompe pasados los 200k)
→ 10 veces más barato
→ 52 veces más rápido que FlashAttention
Si esto es así, van a dejar a Claude obsoleto.
Tocará probarlo en… https://t.co/s7QbxAWn1RIf SubQ's numbers are real, Claude has a serious problem.
SubQ vs Opus 4.6:
→ 12M context tokens (Claude breaks beyond 200k)
→ 10x cheaper
→ 52x faster than FlashAttention
If this is the case, they will make Claude obsolete.
Will have to test it on -
Quadratic Attention Limits Frontier Models After 8 Years
By
–
Attention Is All You Need (2017): most cited ML paper of the decade.
— Sumanth (@Sumanth_077) 5 mai 2026
For 8 years, every frontier model has been built on quadratic attention. Process every possible word-to-word relationship. Compute explodes with context length. Accuracy degrades past 200k tokens.… https://t.co/tSJOH4A2CuAttention Is All You Need (2017): most cited ML paper of the decade. For 8 years, every frontier model has been built on quadratic attention. Process every possible word-to-word relationship. Compute explodes with context length. Accuracy degrades past 200k tokens.
-
World’s First Native Color LiDAR Sensor Unveiled
By
–
World’s first native color LiDAR sensor. I want one of these puppies. https://t.co/Vq44kbtzRO
— Bilawal Sidhu (@bilawalsidhu) 5 mai 2026World’s first native color LiDAR sensor. I want one of these puppies.
-

Anthropic study finds 6% of Claude conversations non-work related
By
–
Anthropic just dropped a study that should make every AI user uncomfortable. They analyzed 1 million Claude conversations from March and April 2026, filtered to roughly 639,000 unique users, and found that about 6% of all conversations weren't about code, work, or homework.
-
Northeastern University Unveils Hybrid Wheel-Leg Robot for Versatile Terrain
By
–
Northeastern University Unveils Hybrid Wheel-Leg #Robot That Rolls Fast and Walks Over Rough Terrain
— Ronald van Loon (@Ronald_vanLoon) 5 mai 2026
by @tweetciiiim
#Robotics #ArtificialIntelligence #Innovation #Technology pic.twitter.com/KfukzRu99FNortheastern University Unveils Hybrid Wheel-Leg #Robot That Rolls Fast and Walks Over Rough Terrain
by @tweetciiiim #Robotics #ArtificialIntelligence #Innovation #Technology -
Naive RAG vs Agentic RAG Visual Comparison Explained
By
–
Naive RAG vs. Agentic RAG, explained visually:
— Akshay 🚀 (@akshay_pachaar) 5 mai 2026
Naive RAG breaks in 3 ways:
↳ It retrieves once and generates once. If the context isn't relevant, the system can't search again.
↳ It treats every query the same. A simple lookup and a multi-hop reasoning task go through the… pic.twitter.com/yJiaYNUPukNaive RAG vs. Agentic RAG, explained visually: Naive RAG breaks in 3 ways: ↳ It retrieves once and generates once. If the context isn't relevant, the system can't search again. ↳ It treats every query the same. A simple lookup and a multi-hop reasoning task go through the
-
Sudo R1: A Robot Trained Entirely in Simulation
By
–
Sudo R1: A #Robot Trained Entirely in Simulation
— Ronald van Loon (@Ronald_vanLoon) 5 mai 2026
by @sudo_robotics#Robotics #ArtificialIntelligence #Innovation #MI #ML #Tech pic.twitter.com/6lnk3UgvCUSudo R1: A #Robot Trained Entirely in Simulation
by @sudo_robotics #Robotics #ArtificialIntelligence #Innovation #MI #ML #Tech