Every time I try Opus 4.1 in the Claude client, it just breaks – goes around in circles even for relatively simple tasks and then throws up an error. I'm sure it is not that Opus is bad, but it seems to have outgrown its Claude interface shackles
LLMS
-

Mistral Medium 3.1 Released
By
–

Mistral released Mistral Medium 3.1! It is now available as a default model on Le Chat as well as in the API as mistral-medium-2508
-

Fixing Transformers GPT-OSS MoE Finetuning Issues
By
–
"i thought the transformers gpt-oss MoE finetuning was broken, how did you get it working?"
-

Claude Sonnet 4 1M Context Window Access Requirements
By
–
Notes on the new 1m context window for Claude Sonnet 4: https://
simonwillison.net/2025/Aug/12/cl
aude-sonnet-4-1m/
… You need to send a beta header of context-1m-2025-08-07 and be on tier 4, which means you have purchased at least $400 in API credits -

LFM2-VL Models: 1.6B and 450M Vision-Language Released
By
–
We also provide an inference and a fine-tuning Colab notebooks. LFM2-VL-1.6B: https://
huggingface.co/LiquidAI/LFM2-
VL-1.6B
… LFM2-VL-450M: https://
huggingface.co/LiquidAI/LFM2-
VL-450M
… -

Liquid Releases Fast 450M and 1.6B Parameter VLM Models
By
–
Liquid just released two 450M and 1.6B param VLMs! They're super fast and leverage SigLIP2 NaFlex encoders to handle native resolutions without distortion. Available today on @huggingface
! -

Mistral Medium 3.1 Released with Performance Improvements
By
–
Introducing Mistral Medium 3.1. Overall performance boost, tone improvement, smarter web searches. Try it now in Le Chat (default model) or via our API (`mistral-medium-2508`).
-
Open-source world models welcome announcement
By
–
Welcome to open-source world models! https://t.co/IXVxZdGm4r
— clem 🤗 (@ClementDelangue) 12 août 2025Welcome to open-source world models!
-
CulturalGround: Extracting More Training Data from Cultural Knowledge Bases
By
–
Importantly, CulturalGround is further evidence that we are not out of data yet and we can squeeze out more and more training data from the knowledge bases. By just rethinking where cultural data resides and systematically generating factual questions and answers about the
-

CulturalGround Fine-tuning Achieves State-of-the-Art Performance on Cultural Benchmarks
By
–
To demonstrate the effectiveness of CulturalGround, we fine-tune an existing multimodel model Pangea on a subset of the dataset. The resulting model achieves state-of-the-art performance for its model size on multiple cultural benchmarks in PangeaBench and other benchmarks. Below