@zephyr_z9 On the topic of Moonshot affording compute, there's a risk… I think the ultra-sparse approach is failing to handle logical reasoning and problem solving like frontier models. Made a new benchmark digging deeper:
LLMS
-

DeepSeek V4 Open Models Lag Behind Frontier AI
By
–
@teortaxesTex About DeepSeek V4 being able to compete with the frontier, I made a new benchmark that suggests open models (particularly the new ultra-sparse ones) are qualitatively worse at problem solving and logical reasoning:
-
Open Weights Model Performance: MoE Sparsity Challenges
By
–
I will add the remaining two (?) models rumoured for release early next week and finalize it with a blog post… In the meantime, if you have any theories why open weights struggle (my theory is that it's MoE/sparsity induced) — let me know!
-
Codex vs Gemini: Frontend Coding Skills and App Integration Race
By
–
Do you think Codex will reach Gemini-level front-end coding skills faster, or will the ability to connect and use Gemini through the Codex App happen first?
-
@testingcatalog — 2026-02-15
By
–
BREAKING 🚨: xAI is working on Parallel Agents mode and Aren Mode for the upcoming Grok Build.
— 🚨 AI News | TestingCatalog (@testingcatalog) 15 février 2026
With Parallel Agents, users will be able to spawn up to 8 coding agents in parallel, while in Arena mode, we will likely see a tournament-style evaluation. pic.twitter.com/324TDKn3PmBREAKING : xAI is working on Parallel Agents mode and Aren Mode for the upcoming Grok Build. With Parallel Agents, users will be able to spawn up to 8 coding agents in parallel, while in Arena mode, we will likely see a tournament-style evaluation.
-

Open Weights Models Lag Behind Frontier on Logical Reasoning
By
–
PREVIEW: The Joy Of Benchmarks (Q1'26) My new #AI benchmark on out-of-domain programming languages (joy) suggests that open weights models are qualitatively *far* behind the frontier on logical reasoning and problem solving… The newest models: GLM-5, Minimax M2.5, and Kimi
-

LLMs Transform Data Cleaning From Manual to Automated Preparation
By
–
Could LLMs finally end the nightmare of manual data cleaning? Researchers from SJTU, Tsinghua, Microsoft Research, MIT, and Alibaba present a comprehensive survey on the future of application-ready data preparation. They detail a massive paradigm shift from rigid, rule-based
-

Book: Unlocking Data with Generative AI and RAG
By
–
New updated 2nd Edition, by @keithbourne v/ @PacktDataML "Unlocking Data with Generative AI and RAG — Learn AI Agent Fundamentals with RAG-powered Memory, Graph-based RAG, and Intelligent Recall" Get the book here: http://
amzn.to/49zsIkb -

LLM Engineer’s Handbook — Build and Train Large Language Models
By
–
LLM Engineer's Handbook — Master the art of engineering Large Language Models #LLMs from concept to production: http://
amzn.to/4dUQrv6 v/ @PacktDataML Implement robust data pipelines and manage LLM training cycles Create your own LLM and refine with the help of hands-on -

RAG as a Control Layer: Beyond Simple Search and GPT
By
–
RAG isn’t “search + GPT”. It’s a control layer:
• limits hallucinations
• enforces evidence
• defines what the model is allowed to know LLMs generate text.
RAG defines truth.
