i.e I wouldn't be at all shocked to hear that there are (were?) OOM inefficiencies in xai's training code base.
@jeremyphoward
-
Jax vs C: inefficient implementation comparison
By
–
IIUC the previous training code at xai was super inefficient. Perhaps the 10x is comparing to that? So it's not really "jax vs C" but "crappy jax impl vs less crappy C impl"?
-
LLMs should aid human learning, creativity, and experimentation
By
–
I feel that the trend towards training models to autonomously go off and try to do everything themselves is anti-human. We should, IMO, be training LLMs to support humans in their learning, creativity, and iterative experimentation.
-
GPT 5.5 improving while Claude models worsening, no clear winner
By
–
GPT 5.5 seems to be improving in that direction now, and Claude models are getting worse at it, so I don't think there's a clear winner now.
-
Need better evaluation of models for human-AI cooperation
By
–
We desperately need better ways of evaluating models. Something that shows how helpful they are at working hand-in-hand with humans to help them get stuff done in a cooperative/iterative way. The Claude models have consistently been better at this, and the market rewards that.
-
Gemini Flash 3.5 criticized for prioritizing evals over user helpfulness
By
–
Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's been trained to max evals, not to be helpful to humans. It goes off and does random crap "for me" rather than just doing what I asked.
-
User critiques Opus for overconfident responses and cost
By
–
I've stopped using Opus for brainstorming/strategizing, because it keeps wanting to jump to a conclusion and the end of every response. It's too confident it knows the answer every time. It makes it hard to have a back-and-forth. Also, it's too expensive vs Codex 5.5 sub.
-
UX pattern for LLM tool failures and dialog integration
By
–
Yes! There's a lot of nice UX approaches this style opens up.
For instance, if the LLM runs a bit of code that's blocked by the sandbox, we don't just show a `y/n/a(ll)` prompt, but instead stop the tool loop and insert the failed code into the dialog. -

Anthropic changes ‘interactive’ billing for Claude and Agent SDK
By
–
This is misleading. This policy redefines the term "interactive" to mean "using an Anthropic front-end". If you use `claude -p` or Agent SDK to do something interactively, it now uses credits, not your subscription limits. So the "interactive use" heading saying "unchanged"