Humanity's Last Exam (HLE) is a rigorous intelligence benchmark featuring over 2500 problems crafted by experts in mathematics, natural sciences, engineering, and humanities. Most models score single-digit accuracy. Grok 4 and Grok 4 Heavy outperform all others.
GENERATIVE AI
-
Grok 4 Unveiled as World’s Smartest AI Model
By
–
We just unveiled Grok 4, the world’s smartest artificial intelligence. 🧵
— xAI (@xai) 11 juillet 2025
Grok 4 outperforms all other models on the ARC-AGI benchmark, scoring 15.9% – nearly double that of the next best model – and establishing itself as the most intelligent AI to date. pic.twitter.com/0PADgAXNpEWe just unveiled Grok 4, the world’s smartest artificial intelligence. Grok 4 outperforms all other models on the ARC-AGI benchmark, scoring 15.9% – nearly double that of the next best model – and establishing itself as the most intelligent AI to date.
-
Man versus Machine: Who Wins at Coding?
By
–
This should be interesting. Who wins at coding. Man or machine?
-
Lovable Stands Out Among AI Coding Tools Like Bolt and Cursor
By
–
back in the day I played with Bolt, Replit, Cursor, and Lovable and I found Lovable was the most… lovable? It works pretty well, I like the interface, and I'm happy with the results I'm getting. And you can sync to Github and then work on your project with Cursor if you want
-
AI Tools Overload: Managing the DM Assault
By
–
thanks for letting me know. I'm getting hammered by all these tools in the DMs and it just feels like an assault.
-
How Frontier AI Models Are Built: The Iterative Process
By
–
How frontier AI models are built 1. identify task model cannot solve
2. create task eval
3. collect new *eval-specific* training data
4. train model
5. model can do the task now
6. if not (AGI achieved): goto step 1 GLUE, MMLU, MATH, AIME, HLE, GPQA, ARC a tale as old as time -
AI Assistant Functions Like Professional Designer
By
–
Depends on your request, it works as a real designer would.
-
Vision VAE as Ultimate Tokenizer: Unicode Efficiency
By
–
Very cool work direction but also fair question.
I wonder if ultimately is a little vision patch VAE the ultimate "tokenizer"? Unicode + UTF-8 is just too high description length. -
Are People Unable to Detect AI Quality Anymore?
By
–
Maybe people are just not good at detecting quality at this point? (Based on my dive into the actual answers people prefer, certainly plausible)
-
LM Arena Benchmark Relevance Decline Among AI Makers
By
–
It is weird that leading LM Arena went from being the big benchmark every AI maker was aiming for to being not mentioned much in recent releases. Post-Llama 4 reputation hit? Post GPT-4o Sycophantic Apocalypse realization that arena scores were easily optimized? Temporary blip?
