constant rumors of GPT5-reasoning, GPT5-mini and GPT5-nano coming in "early August" now. according to this @tomwarren piece, o3-mini-level open model launch also coming by next week.
LLMS
-

Claude 4 Agent Detects 7 of 10 Implanted Concerning Behaviors
By
–
Our third agent was developed for the Claude 4 alignment assessment. It red-teams LLMs for concerning behaviors by having hundreds of probing conversations in parallel. We find the agent uncovers 7/10 behaviors implanted into test models.
-

RLVR Models Sacrifice Diversity for Base Model Preferences
By
–
RLVR'd model would fail to answer questions that its base model previously could. "The Invisible Leash: Why RLVR May Not Escape Its Origin" This research shows that RLVR just boosts the base model’s favorite answers (raising pass@1) and would sacrifice diversity.
-
Spark API Endpoints for GPT-4o LLM Calls Without Proxying
By
–
Not in its current state, no – thankfully you don't need to proxy API keys for LLM calls with Spark because they run an API endpoint for you (just for gpt-4o and gpt-4o-mini at the moment though)
-
From iOS Development to LLM Personality Design: Tech Evolution
By
–
Working in tech then: how to build good ios apps on 3G
Working in tech now: how to give LLMs personalities -

Managed MCP Servers: Secure LLM Agent Tools with Governance
By
–
Equipping LLM agents with tools shouldn’t mean losing governance or security. Model Context Protocol (MCP) has taken off in recent months, and we’re excited to introduce managed MCP servers with Mosaic AI and Unity Catalog integration. Now your your AI models can securely
-
Free LLM Tool Available Now at Datasette
By
–
I'm currently giving it away for free https://
llm.datasette.io -
Claude Sonnet 4 Custom Prompt Project Implementation
By
–
Claude Sonnet 4 with a custom prompt in a project
-

Kimi K2: Open-Source Trillion-Parameter MoE Model Released
By
–
Kimi K2 = Hype or Not?
— Louis-François Bouchard 🎥🤖 (@Whats_AI) 24 juillet 2025
It is the first open‑source trillion‑parameter MoE model
• 1 T total / 32 B active parameters → giant yet efficient
• Outscores GPT‑4.1 & Claude‑4 (65.8 % SWE‑Bench, 53.7 % LiveCodeBench)
• 128 k context, MIT‑licensed, no vision implemented…… pic.twitter.com/tiouZDLCScKimi K2 = Hype or Not? It is the first open‑source trillion‑parameter MoE model • 1 T total / 32 B active parameters → giant yet efficient
• Outscores GPT‑4.1 & Claude‑4 (65.8 % SWE‑Bench, 53.7 % LiveCodeBench)
• 128 k context, MIT‑licensed, no vision implemented… -

Gemini 2.5 Pro Achieves Gold-Winning IMO Performance
By
–
Gemini 2.5 Pro Capable of Winning Gold at IMO 2025 Yichen Huang, Lin F. Yang: https://
arxiv.org/abs/2507.15855
v2
… #ArtificialIntelligence #DeepLearning #MachineLearning
