Flops on the sestina in a very weird way. It seems absolutely unable to hit the end word scheme that many smaller models succeed at (albeit without the s constraint). Have tried it a few times in different ways. Curious.
@emollick
-

Grok 4 Passes Lem Test with Most Coherent Narrative
By
–
Grok 4 passes the Lem test first try, with the most coherent narrative yet.
-
AI Model Lacks Safety Documentation and Transparency Measures
By
–
Impressive model based on a few minutes of playing, but disappointing to see no mention at all of a model card, red teaming, yesterday's incident, or how they are going to address the process issues they keep having.
-

Grok 4 Successfully Creates Shader Without Errors
By
–
Grok 4 creating the shader (no errors). https://t.co/zoteK6KGYr pic.twitter.com/YeWrL2Yb5u
— Ethan Mollick (@emollick) 10 juillet 2025Grok 4 creating the shader (no errors).
-
Scale, Tool Use, and Multimodal: AI Development Path Forward
By
–
It looks like scale + tool use + multimodal remains the chosen path forward.
-
Regrettably Immanentizing the Eschaton: AI and Existential Risk
By
–
Really leaning into regretfully immanentizing the eschaton
-
Grok 4 reaches 10^27 FLOPs, outperforms Gemini 2.5
By
–
Looks like Grok 4 is 10^27 FLOPs given their graphs? HLE score is 26% without tools, Gemini 2.5 is 21.6% without tools. Curious what the tool piece is.
-
Demonstrating Advanced AI: The Challenge of Showcasing Grok 4
By
–
Among other things with the Grok 4 launch, it will be interesting to see how you demo a (presumably) very smart model. We are getting to the point where current AIs already do a lot of impressive things, so it is harder and harder to show to non-experts what a new model does.
