Since Anthropic publish their system prompts we can generate a diff between Claude Opus 4.6 and 4.7 – here are my notes on what's changed
LLMS
-

Benchmarking quantization levels with TurboQuant for best context fit
By
–
Tried more quants (4-bit, 5-bit) on upstream llama.cpp Next up is benchmarking TurboQuant for the same quants The goal is finding the best quant that fits with the highest context in q8_0/turbo3 asymmetric
-
Nous Research Hermes outperforms as preferred language model
By
–
I use @NousResearch
’s Hermes. It is better -
Who Trains Llama 70B Today Insufficient Workload
By
–
who trains llama 70B these days though. too much of a toy workload
-

Qwen3.5 9B benchmarked: TurboQuant avoids OOM at full context
By
–
Benchmarked the same Qwen3.5 9B UD-IQ3_XXS GGUF on an RTX 3070 8GB using – Latest upstream llama.cpp VS – TheTom's TurboQuant llama.cpp fork TurboQuant allowed me to reach full context length without OOM, more in the screenshot below
-
Grok AI receives criticism for low quality outputs
By
–
grok is great for those who don't care what they get
-
OpenAI leads reasoning capabilities, DeepSeek close behind
By
–
I feel like OpenAI is the only lab that really nailed reasoning. DeepSeek was probably the closest, but need to see a frontier model from them to be sure. Gemini's reasoning was quite weird and all over the place. Claude's reasoning never used to matter (non thinking were
-
NVIDIA Vera Rubin outperforms Mac Mini for Hermes model
By
–
It is an NVIDIA Vera Rubin. Can run Hermes better than a Mac Mini. Like thousands of times faster. 🙂
-

Master Any LLM: Comprehensive Guide to Large Language Models
By
–
Master Any #LLM
by @ingliguori #GenerativeAI #ArtificialIntelligence #MachineLearning #MI