A few quick observations on Grok 4:
1) Hidden CoT with very little information in the reasoning trace
2) Uses web search a lot (not just searching X)
3) Have not seen it use code to run calculations or solve non-coding problems yet, generally less aggressive about tools than o3
LLMS
-
Grok 4 Analysis: Hidden CoT, Web Search Integration
By
–
-
Grok 4 Scores 50.5% on HLE Benchmark, Score Still Rising
By
–
Pour l’AGI, c’est le HLE qui compte avant tout c'est un peu la mesure de l'AGI. Et d’après ce qu’on m’a dit, le score n’est même pas encore stabilisé ! Le dernier benchmark HLE passé par GROK 4 affiche 50,5 %. Plus ils le laissent tourner, plus le score augmente. On ne connaît
-

Grok 4 Achieves Perfect AIME 2025 Score and Doubles ARC-AGI Performance
By
–


Alors là… pour l'instant GROK 4 sur le papier, grosse percée ! Sur le benchmark AIME 2025 il a fini le Game (100% de score). Premier modelé à réussir. Sur l'ARC-AGI, un x2 par rapport à ce qu'il se fait aujourd'hui. Et surtout sur le HLE (humanity last exam), il fait
-
Grok 4 for Reasoning, Flux Kontext for Photo Restoration
By
–
Oui il est pas fait pour ça grok 4, c'est surtout pour son raisonnement Pour les photos je te conseille flux kontext, j'en ai restauré quelques une encore hier. Apres tout dépends de l'état de la photo…
-
xAI Grok 4 RonnaFLOP model performance benchmarks scaling
By
–
I suspect the next few weeks after Grok 4 follows the same pattern as Grok 3 xAI beats everyone to market with the first RonnaFLOP model. The benchmarks show the 10-20% improvement the scaling law suggests. In the coming months, the other labs release their RonnaFLOPs, catch up.
-
Convert GitHub repos to LLM-optimized prompts easily
By
–
Pro tip: take any github repo url, change the “g” to a “u” (like “uithub”) and you’ll have a copyable, LLM-optimized prompt that contains a structured version of the repo!
-

Claude 4 Performance Comparison Through Processing Pipeline
By
–
For comparison, here’s Claude 4 going through the same pipeline. pic.twitter.com/E4xeJeedGp
— Pietro Schirano (@skirano) 10 juillet 2025For comparison, here’s Claude 4 going through the same pipeline.
-

AI Model Fails at Sestina Poetry Generation Task
By
–
Flops on the sestina in a very weird way. It seems absolutely unable to hit the end word scheme that many smaller models succeed at (albeit without the s constraint). Have tried it a few times in different ways. Curious.
-

Grok 4 vs Claude 4: Design Capabilities Comparison
By
–
Grok 4 is mostly okay at designing…
— Pietro Schirano (@skirano) 10 juillet 2025
Claude 4 still comes out on top here. pic.twitter.com/2bssbIgPgCGrok 4 is mostly okay at designing… Claude 4 still comes out on top here.
-

Grok 4 Passes Lem Test with Most Coherent Narrative
By
–
Grok 4 passes the Lem test first try, with the most coherent narrative yet.