These new open source models (GLM, Kimi) continue to be odd. Great stats, some solid performances, but also fail tests that DeepSeek & smaller closed models have beaten for months.
LLMS
-
Kimi K2 Model Exhibits Significant Hallucination Issues
By
–
Kimi K2 really hallucinates a lot in my limited testing so far, and is very happy to make up new details if it judges that it would improve the punch of a paragraph.
-

Alibaba Qwen3 Models Released: Thinking and Translation Capabilities
By
–
Alibaba's Qwen team dropped two more models: —A Qwen3-Thinking update, now competitive with Gemini 2.5 Pro, o4-mini, and DeepSeek R1 across multiple benchmarks
—Qwen3-MT, an AI translation model with support for 92+ languages -

Meta Hires OpenAI Researcher Shengjia Zhao as Chief Scientist
By
–
Meta Superintelligence Labs finally found its chief scientist: ex-OpenAI researcher Shengjia Zhao Zhao helped create OpenAI's GPT-4, o1, o3, 4.1, and mini models, and will set MSL’s research direction alongside chief AI officer Alexandr Wang
-
Model 4.5 Excels at Rewriting While Preserving Original Style
By
–
4.5 is the only model I trust with writing, esp re-writing without changing the style – all others don't understand the task
-

LLM Arena Models Proliferation: Concerns About Score Gaming
By
–
Suddenly there are tons more weird LLM arena models – cuttlefish, kraken, etc. I just hope we are not going to see a repeat of the Llama 4 incident, where different versions of the same model are being tuned to max out the arena score
-
When Will AI Model Sizes Reach Avogadro Scale?
By
–
How long before model sizes are measured in Avogadros (multiples of 6 x 10^23)?
-
Using Claude and ChatGPT for Deep Topic Recommendations
By
–
really depends what topic you want to go deep on. claude and chatgpt are pretty good at recommending, particularly if you say "i read x, I want more on a similar topic"
-

OpenAI vs Google: LMArena Rankings Competition Analysis
By
–
As we get ready for GPT-5, it's useful to look back at how often labs featured in the Top 5 of @lmarena_ai over the last 1.5 years. The competition is primarily between OpenAI and Google. Average appearances overall and specifically in 2025:
– OpenAI: Overall: 2.2; in 2025: 1.7 -
Hallucinations in RNNs: Unreasonable Effectiveness Revisited
By
–
I believe this is true, I used the word in my “Unreasonable Effectiveness of RNNs” post from 2015, and as far as I can remember I also hallucinated it.
