Actually, I don’t think HLE is a great measure of usefulness. We’re moving away from these benchmarks in favor of making Grok maximally useful for actual engineering.
CODE
-
Codex 5.3 Release Generates Strong Momentum in Developer Community
By
–
codex momentum is strong, and many people are feeling just how big of a leap 5.3 is. if your organization hasn't tried codex yet, it's worth revisiting.
-

Minimax M2.5 Launch: Advanced AI for Productivity Workflows
By
–
Minimax M2.5 is live on Poe! The new Minimax M2.5 is designed for real world productivity, with a focus on planning‑driven, multi‑step execution across complex digital workflows. It excels at coding, cross‑tool context switching, and agent style task coordination. Also, it
-
Elon’s AI behind in coding; OpenAI builds decentralized computer products
By
–
Elon's AI is behind, particularly for coding. Also, OpenAI has a whole product team that's building a variety of products that could run a decentralized computer. Elon has a bit of a reputation, and so does Sam. So, God knows, maybe it was just the pitch. Steve Jobs bought Siri
-
Use unfiltered lists and create your own AI for X
By
–
Just use lists. Lists are not filtered. Write your own AI to read X. That's what I'm doing.
-

Open-source multimodal RAG framework discovery and integration
By
–
Ever since I started working on Memory, I've been seeing RAG products every day. Discovered a one-stop RAG framework, open-source, MIT licensed. This library can be considered a multimodal superset based on LightRAG. Building on the LightRAG architecture, it provides a
-

Open Models Overfitting Benchmarks While Losing Reasoning Ability
By
–
@xeophon On the topic of swe-rebench and lower scores, another data point for you: my own analysis suggests open models are overfitting to popular patterns/benchmarks while failing to get better at logical reasoning / problem solving:
-
Codex vs Gemini: Frontend Coding Skills and App Integration Race
By
–
Do you think Codex will reach Gemini-level front-end coding skills faster, or will the ability to connect and use Gemini through the Codex App happen first?
-
@testingcatalog — 2026-02-15
By
–
BREAKING 🚨: xAI is working on Parallel Agents mode and Aren Mode for the upcoming Grok Build.
— 🚨 AI News | TestingCatalog (@testingcatalog) 15 février 2026
With Parallel Agents, users will be able to spawn up to 8 coding agents in parallel, while in Arena mode, we will likely see a tournament-style evaluation. pic.twitter.com/324TDKn3PmBREAKING : xAI is working on Parallel Agents mode and Aren Mode for the upcoming Grok Build. With Parallel Agents, users will be able to spawn up to 8 coding agents in parallel, while in Arena mode, we will likely see a tournament-style evaluation.
-
AI: a tool that replaces tasks, not skills
By
–
It can't be said enough but: AI is first and foremost a tool. It does not replace skills. It replaces tasks. AI will not take the job of a good developer. And an average person, just with AI, does not magically become a good developer.