BREAKING : Mistral AI released Devstral Small and Medium 2507, topping the scores on the SWE bench among open models. – devstral-small-2507 – $0.1/M input tokens and $0.3/M output tokens.
– devstral-medium-2507 – $0.4/M input tokens and $2/M output tokens.
@testingcatalog
-

Mistral AI launches Devstral Small and Medium 2507
By
–
-

Gemini 3.0 Pro reference discovered in Gemini-CLI commit
By
–

BREAKING : Gemini 3 reference has been spotted in the Gemini-CLI commit! gemini-beta-3.0-pro
-

OpenAI to Release New Open-Weight Model Matching o3-mini Performance
By
–

BREAKING : According to The Verge, OpenAI will release its open weight model next week and it will be performing at the o3-mini level.
-

Claude Desktop App May Add MCP Connectors Directory
By
–



Claude desktop app may get a Connectors Directory, a curated list of remote and desktop-specific MCPs. Lowering an entry barrier into MCPs
-
Vidu Q1 Model Adds Reference-to-Video Tool Support
By
–
Now you can use Vidu Q1 with the Reference-to-Video tool! It supports up to 7 different references (images) to be used with the latest Q1 model.
— 🚨 AI News | TestingCatalog (@testingcatalog) 8 juillet 2025
There, you can simply blend all your characters and scenes into different 5-second videos with HD quality. https://t.co/Te1wWz0PLh pic.twitter.com/4PoDMEj4qfNow you can use Vidu Q1 with the Reference-to-Video tool! It supports up to 7 different references (images) to be used with the latest Q1 model. There, you can simply blend all your characters and scenes into different 5-second videos with HD quality.
-
Hands-on review of Proactor AI assistant features
By
–
I got a chance to test Proactor AI during the weekend 👀
— 🚨 AI News | TestingCatalog (@testingcatalog) 8 juillet 2025
It is one of the few AI tools that acts proactively and works as your real assistant. It can transcribe meeting recordings on the fly, generate suggestions, organise to-do lists, assist with execution and a lot more! https://t.co/HruD8o4VpV pic.twitter.com/Hfst49GgEsI got a chance to test Proactor AI during the weekend It is one of the few AI tools that acts proactively and works as your real assistant. It can transcribe meeting recordings on the fly, generate suggestions, organise to-do lists, assist with execution and a lot more!
-

Claude Neptune v3 shows competitive math performance against top-tier models
By
–

BREAKING : Some users who have received access to "Claude Neptune v3" are reporting that it can consistently solve math problems at a level of o3 Pro and "Kingfall". The next leap? h/t @No_name_890098
-

Grok 4 Benchmarks vs Other Models
By
–


Grok 4 early benchmarks in comparison to other models. Humanity last exam diff is Visualised by @marczierer
-

Grok 4 SOTA avec des scores élevés
By
–


BREAKING : Grok 4 will be SOTA – 35% on HLE, 45% with reasoning
– 87-88% on GPQA – 72-75% on SWE Bench (for Grok 4 Code) * Not official benchmarks, I plotted these based on the leaked scores and results from the web for other models -

Claude Neptune V3 Accessed by Red Teams
By
–
BREAKING : Red teams are getting access to "claude-neptune-v3" "Neptune" was initially introduced as a new safety system shortly before the Claude 4 launch. Is it a Claude 4.5 time?