BullshitBench: GPT-5.5 and 5.5-Pro update! They did NOT do well – 5.5 about the same level as GPT-5.4 (around 30-35 rank, 45% pushback). GPT-5.5-Pro did WORSE – only about 35% pushback. I must say the Pro result kind of shocked me. This is actually interesting, what this tells
LLMS
-

DeepSeek V4 Pro Flash Day-0 vLLM Support Long-Context
By
–
Day-0 support for @deepseek_ai V4 Pro and Flash on vLLM — a new generation of DeepSeek model, purpose-built for tasks up to 1M tokens. Alongside the release, we're publishing a first-principles walkthrough of the new long-context attention and how we implemented it in vLLM. x.com/deepseek_ai/st…
-

DeepSeek V4 Launch with SGLang Optimizations and RL Pipeline
By
–
DeepSeek V4 by @deepseek_ai just dropped! SGLang is ready on Day 0 with a full stack of optimizations from architectures to low-level kernels. We also deliver a verified RL training pipeline in Miles (by @radixark) for V4 at launch: Native "ShadowRadix" Design: DeepSeek V4's
-
Day-One Model Support Commitment for Major Releases
By
–
"Day-one model support: every major release, compatible at launch" <- this is quite a commitment! Awesome news! Congrats!
-
DeepSeek-V4 Million-Token Context LLM Agentic Workflows
By
–
DeepSeek-V4 is here — a million-token context, 1.6T parameter powerhouse optimized for agentic workflows. Out of the box, on DeepSeek-V4-Pro, NVIDIA Blackwell Ultra delivers over 150 TPS/user interactivity for agentic workflows. And we’re just getting started. Expect these
-
GPT 5.5 Thinking Launches on ChatLLM with 50% Discount
By
–
GPT 5.5 Thinking Is Rolling Out On ChatLLM On Abacus AI We also have a very exciting promotional offer!! Try it, at 50% off for a short period of time
-
Hugging Face Models Directory Should Support Quantization Filtering
By
–
Ideal fix would be for the HF models directory to grow a direct understanding of the structure of those kinds of repos and treat them as individual models that can be listed separately, including filter by quantization type
-
Naming failure modes matters in nondeterministic AI systems
By
–
Agree on the pattern, disagree on it being pointless… putting names on the failure modes matters more when the system is this nondeterministic.
-
Medical imaging AI proves effective but LLMs lack real world evidence
By
–
Superhuman interpretation AI for medical images, such as mammography and endoscopy, has been proven to improve diagnostic accuracy in multiple randomized trials but mostly not implemented. But LLMs for clinical decision support have little real world medicine proof, but are
-
Chronicle Enhances AI Agent Memory and Context Understanding
By
–
chronicle gives codex rich context and recent memory over what you’re doing: https://
x.com/OpenAIDevs/sta
tus/20462882437680826994/
…
