Also, I don’t know how OP is getting Qwen3.5 27B @ 30 tps on DGX Spark That number is impossible for a Dense model on DGX Spark’s Unified Memory (273 Gbps) My personal experiments showed it’s 4 tokens/sec for that model on the Spark
LLMS
-

MiniMax M2.5 Lightning Attention Architecture for Long Context Scaling
By
–
Lightning Attention architecture used in MiniMax M2.5 is really interesting. The structure is 7 Lightning Attention layers for every 1 traditional SoftMax attention layer, which lets it scale to long contexts while keeping the quality you'd expect from standard transformers. I… pic.twitter.com/HnfTF0W6f1
— Akshay 🚀 (@akshay_pachaar) 14 mars 2026Lightning Attention architecture used in MiniMax M2.5 is really interesting. The structure is 7 Lightning Attention layers for every 1 traditional SoftMax attention layer, which lets it scale to long contexts while keeping the quality you'd expect from standard transformers. I
-
Unified Memory Faster for Loading Large MoE Models
By
–
the issue is that unified memory would still be faster for loading MoEs that are larger than the largest single GPU in terms of Memory
-

RTX PRO 6000 GPU Inference Faster Than Unified Memory After Loading
By
–
This will probably be great for Large single GPUs (e.g. RTX PRO 6000) You’re limited to 40Gpbs initially (during model loading) but then once the model is fully loaded on the GPU it should be extremely faster than Unified Memory speeds for inference
-
Sam Altman Predicts Breakthrough Beyond Transformers Coming Soon
By
–
Ton grave, pose sérieuse.
— VISION IA (@vision_ia) 14 mars 2026
Sam Altman dit qu’une autre percée au-delà des transformers va bientôt arriver, et que les modèles sont désormais assez intelligents pour aider à la découvrir.
L’IA crée une énorme opportunité de reconstruire des catégories entières de produits et de… pic.twitter.com/wTEt3aWv0NGrave tone, serious pose. Sam Altman says that another breakthrough beyond transformers is coming soon, and that models are now smart enough to help discover it. AI creates a huge opportunity to rebuild entire categories of products and make new things possible. A new science,
-
Prompt Engineering’s Persistence: A Sign We’re Far From AGI
By
–
The persisting importance of prompt engineering — and now harness engineering — is one of the best indicators of how far we are from AGI. A general system doesn't need a task-specific harness. And when provided with instructions, it is robust to phrasing variations.
-

Claude Opus 4.6 and Sonnet 4.6 context window expansion
By
–
Claude Opus 4.6 and Sonnet 4.6 now run a full 1 million token context window at standard pricing.
No premium multiplier. No beta header. A 900K-token request costs the same per-token as a 9K one. What this actually means:
→ Load an entire codebase, hundreds of contracts, or a -
ShinkaEvolve: LLMs with Evolutionary Algorithms for Scientific Discovery
By
–
Robert Lange @RobertTLange from @SakanaAILabs on ShinkaEvolve — an open-source framework combining LLMs with evolutionary algorithms for scientific discovery, with insane sample efficiency.
— Machine Learning Street Talk (@MLStreetTalk) 14 mars 2026
His thesis that current systems optimise solutions to fixed problems. Going forwards –… pic.twitter.com/PPqVs5AeZcRobert Lange @RobertTLange from @SakanaAILabs on ShinkaEvolve — an open-source framework combining LLMs with evolutionary algorithms for scientific discovery, with insane sample efficiency. His thesis that current systems optimise solutions to fixed problems. Going forwards — real scientific discovery requires co-evolving the actual problems. By the way – NVIDIA GTC is coming and will showcase breakthroughs in physical AI, AI factories, agentic AI, and inference. Register for virtual GTC for free using my link: nvda.ws/4qQ0LMg and enter raffle to win a DGX Spark 😈
→ View original post on X — @_yutaroyamada, 2026-03-14 05:34 UTC
-

LLMs Transform Reinforcement Learning in Recommendation Systems
By
–
Can recommendation systems truly understand and adapt to your evolving tastes? A team from University of Science and Technology of China, Kuaishou , and others is mapping out how Large Language Models (LLMs) revolutionize Reinforcement Learning (RL) in recommendation systems.
-
Closed Models Gaining Ground Over Open Source Models
By
–
As of recent, the gap between closed models and open models is widening.