exllamav3 is so underrated btw
LLMS
-
Symbolic Verifiers vs LLM Judges for Verification Tasks
By
–
What do you mean with set of prompts? A LLM judge with a rubric? (That’s a different approach; covered in the appendix).
The reason for a symbolic verifier is that it is symbolic. There is no nondeterminism, bias, etc. Doesn’t work for everything, but eg for math there is no -
Frequency of car brand citations in AI EV recommendations
By
–
Hello, the number of times car brands are cited when asking for EV recommendations to different AI models.
-
Hume AI Octave 2 Multilingual Model Announced
By
–
BREAKING 🚨: Hume AI is preparing to release Octave 2 Multilingual model! Here is a sample dialogue between a Robot and a Russian hacker.
— 🚨 AI News | TestingCatalog (@testingcatalog) 29 septembre 2025
"Expressive, natural-sounding voices in 10+ languages, low latency, perfect for real-time transition and conversational use cases" pic.twitter.com/IUIZ8gb8WUBREAKING : Hume AI is preparing to release Octave 2 Multilingual model! Here is a sample dialogue between a Robot and a Russian hacker. "Expressive, natural-sounding voices in 10+ languages, low latency, perfect for real-time transition and conversational use cases"
-
DeepSeek Achieves 50x Attention Efficiency Breakthrough
By
–
DeepSeek casually unlocked 50x attention efficiency in ~1 year > MLA is ~5.6x faster than MHA
> DSA is 9x faster than MLA never doubted you, you big beautiful whale -

What Should Replace MMLU for AI Model Evaluation?
By
–
Final update! MMLU is saturated and has become (rightfully) less popular. What should replace it? – Other knowledge evals like GPQA, MMLU-Pro
– Code evals like LiveCodeBench
– Agentic evals like BFCL
– Other? -

DeepSeek-V3.2-Exp: 50% Cheaper, Better Search
By
–



DeepSeek released DeepSeek-V3.2-Exp build on top of previously released V3.1-Terminus model. It is 50% cheaper and slightly better at search benchmarks. Deep dumping
-
Elon Musk Claims Grok 4 Surpasses PhDs, May Discover Physics Laws
By
–
Elon Musk : Je m’attends à ce que Grok découvre de nouvelles lois de la physique d’ici l’année prochaine.
— VISION IA (@vision_ia) 29 septembre 2025
« Je veux vraiment insister sur ce point. Sur les questions académiques, Grok 4 est meilleur qu’un docteur en chaque matière. Sans exception.
Je m’attends à ce que Grok… pic.twitter.com/zyUjIWezGJElon Musk : Je m’attends à ce que Grok découvre de nouvelles lois de la physique d’ici l’année prochaine. « Je veux vraiment insister sur ce point. Sur les questions académiques, Grok 4 est meilleur qu’un docteur en chaque matière. Sans exception. Je m’attends à ce que Grok
-

Grok 4 Heavy: Advanced AI Model for Complex Problem Solving
By
–
Grok 4 Heavy for the toughest problems. And it gets better almost every week.
