Grok 4 suggests that scaling still works (with the diminishing returns predicted by the scaling law), and that tool use can unlock performance gains. Kimi suggests there continues to be big opportunities from improvements in methods (Muon, etc.). Lots of paths for AI right now.
LLMS
-
Transformers Limited: AI Solved Problems Decades Ago
By
–
Yet another problem AI solved 50 years ago and transformers can't, and yet another onlooker thinking AI is just transformers. https://t.co/BVD4tJICPL
— Pedro Domingos (@pmddomingos) 13 juillet 2025Yet another problem AI solved 50 years ago and transformers can't, and yet another onlooker thinking AI is just transformers.
-
Grok Gemini Claude Cursor AI Tools Comparison Guide
By
–
for me > grok for real time information carving
> gemini for research
> claude for coding and technical tasks
> cursor for dev and coding tasks -
AI Tools Evolution: ChatGPT to Grok Multi-Platform Journey
By
–
My AI tools use journey: chatgpt > stability ai > gemini > hugging face models chatbot > all in one chatbots > claude > grok (sometimes) > claude > gemini > all at once and now grok+claude+gemini
-
WebGPT thinking sections and AI model development history
By
–
For sure, there's a long history that everyone is building on, although I don't recall webgpt having a "thinking" section.
-
Hyperstition via search complicates LLM pre-release testing
By
–
If true, such “hyperstition via search” poses a significant complication to pre-release testing of modern LLMs: xAI could not have plausibly noticed this specific “Hitler” response before Grok’s release, as the Grok 3 “MechaHitler” incident causing it had not yet occurred.
-
Search-enabled LLMs showing hyperstition feedback loops
By
–
Speculatively, this behavior seems to demonstrate accelerated “hyperstition” feedback loops in search-enabled LLMs. That is, Grok appears to be influenced by its own past mistakes, via media reporting, without ever being literally trained on them (via model-weight updates).
-

Grok 4 vs Grok 4 Heavy: differing search-based behavior
By
–
The “Thoughts” from Grok 4’s response (unavailable for Grok 4 Heavy) suggest an obvious explanation for Grok’s behavior—Grok searches, finding news of the recent “MechaHitler” incident. Why Grok 4 rejects this candidate answer, while Grok 4 Heavy does not, is unclear.
-

Young Frontier Labs Dominate LLM Rankings in Days
By
–
two frontier labs just emerged and took top spot in closed and open source llms last week back to back in 2 consecutive days. both super young. today is @xai
’s second birthday btw, and moonshot is 4 months older if this has not significantly updated you on what actual moats -
Using Multiple LLMs for Research and Content Generation
By
–
nah, i used a mix of LLMs to research and generate, then mixed and refined within the chat.