Humanity's Last Exam (HLE) is a rigorous intelligence benchmark featuring over 2500 problems crafted by experts in mathematics, natural sciences, engineering, and humanities. Most models score single-digit accuracy. Grok 4 and Grok 4 Heavy outperform all others.
LLMS
-
Grok 4 Unveiled as World’s Smartest AI Model
By
–
We just unveiled Grok 4, the world’s smartest artificial intelligence. 🧵
— xAI (@xai) 11 juillet 2025
Grok 4 outperforms all other models on the ARC-AGI benchmark, scoring 15.9% – nearly double that of the next best model – and establishing itself as the most intelligent AI to date. pic.twitter.com/0PADgAXNpEWe just unveiled Grok 4, the world’s smartest artificial intelligence. Grok 4 outperforms all other models on the ARC-AGI benchmark, scoring 15.9% – nearly double that of the next best model – and establishing itself as the most intelligent AI to date.
-
xAI launches connectors for Grok with Notion, Slack, etc.
By
–
BREAKING 🚨: xAI is gearing up to release connectors with external apps like Notion, Slack, Gmail and Google Calendar.
— 🚨 AI News | TestingCatalog (@testingcatalog) 11 juillet 2025
"Tools let Grok connect to external services for practical tasks, like chatting in Slack channels, managing emails, or handling calendars."
Full docs below 👀 pic.twitter.com/Dtval0Oo7BBREAKING : xAI is gearing up to release connectors with external apps like Notion, Slack, Gmail and Google Calendar. "Tools let Grok connect to external services for practical tasks, like chatting in Slack channels, managing emails, or handling calendars." Full docs below
-

xAI’s Smart Listening Feature for Grok Voice Mode Announced
By
–
BREAKING : xAI is working on a Smart Listening feature for Grok Voice Mode! It is yet unclear what it does though. “Respond selectively based on your commands” What should I command her?
-
How Frontier AI Models Are Built: The Iterative Process
By
–
How frontier AI models are built 1. identify task model cannot solve
2. create task eval
3. collect new *eval-specific* training data
4. train model
5. model can do the task now
6. if not (AGI achieved): goto step 1 GLUE, MMLU, MATH, AIME, HLE, GPQA, ARC a tale as old as time -
Vision VAE as Ultimate Tokenizer: Unicode Efficiency
By
–
Very cool work direction but also fair question.
I wonder if ultimately is a little vision patch VAE the ultimate "tokenizer"? Unicode + UTF-8 is just too high description length. -
LM Arena Benchmark Relevance Decline Among AI Makers
By
–
It is weird that leading LM Arena went from being the big benchmark every AI maker was aiming for to being not mentioned much in recent releases. Post-Llama 4 reputation hit? Post GPT-4o Sycophantic Apocalypse realization that arena scores were easily optimized? Temporary blip?
-

Open Source Non-Reasoning Model Achieves Competitive Performance
By
–
Nuevo super modelo de programación no-razonador, que logra unos números muy competitivos colocándose como el mejor modelo open source de este tipo! Eso sí, 1 trillion de parámetros así que olvidaos de ejecutarlo en local fácilmente. Aún así un regalo brutal para la comunidad
-
Lex Fridman Praises Grok 4 and xAI Team Achievement
By
–
Grok 4 is super strong. Congrats to Elon and xAI team, they cooked
-

Prompt Engineering Evolution: Context Engineering and Intent Encoding in 2025
By
–
i think Prompt Engineering is "growing up" in two different ways in 2025. the first is "Context Engineering", which is well covered by @dexhorthy
's talk (now >100k views!) but Prompts = Intent + Context and as @sgrove points out the task of encoding intents/goals/principles
