We partnered with @FireworksAI_HQ to build an efficient traceability judge. We fine-tuned an @Alibaba_Qwen model to detect 'perceived errors' on each production trace. It matched or exceeded the performance of state-of-the-art models and
LLMS
-

AI agents fail at prototype stage, real value is usability
By
–
AI agents fail when they stop at the prototype stage • Multimodal input
• Structured output
• APIs & interfaces
• Tools + memory + RAG Real value comes from usability, not capability alone. Via Giuliano Liguori (
@ingliguori
) #AI #AIAgents #Tech -
Kimi k2.7 and Qwen 3.8: huge victory and excitements to come
By
–
My opinion: Kimi k2.7 and/or Qwen 3.8 A huge victory for all of us and really excited for the release! Plus: remember that closed-source releases are also on the horizon: Sonnet 5, ChatGPT 5.6 and many more Super ultra excited for the week (or the
-

Sakana AI releases Fugu, a model API for autonomous orchestration
By
–
/4 Sakana AI just released Fugu, a new model API built around autonomous model orchestration. The stronger version, Fugu Ultra, is built for harder multi-step tasks. It scores 73.7 on SWE Bench Pro, 82.1 on TerminalBench 2.1, 93.2 on LiveCodeBench, 50.0 on Humanity’s Last Exam,
-

GLM-5.2 open-weight model with 1M context window released
By
–
/3 http://
Z.ai drops GLM-5.2 weights on Hugging Face. GLM-5.2 is a flagship open-weight model built for long-horizon tasks, especially coding and agentic work. Its biggest headline is a stable 1M-token context window, giving it room to handle large codebases and -

MiniMax M3: Open Weights with Multimodal & Context
By
–
/2 MiniMax opens the model weights of Minimax M3 to the public. M3 is the first open-weight model to combine frontier coding, a 1 million token context window, and native multimodal support (images and video) in one package. That combo was locked behind paid APIs until now. The
-

Invitation to LangChain Meetup Berlin with DB and Zalando on July 16
By
–
Berlin You are invited to a LangChain Meetup with DB Engineering & Consulting, and @Zalando on July 16. Confirm your presence to reserve your spot
https://luma.com/e6onqqkz -
Chinese models efficiency with MoE and open source vs proprietary AI
By
–
Is there also not a driver in Chinese models have more efficiency w/ MOE and other innovations in open source stack vs proprietary AI models? And also maybe because US players are providing data centres to most of the world?
-

Add the gemini-interactions-api skill with npx
By
–
—
For your agents: > npx skills add google-gemini/gemini-skills –skill gemini-interactions-api –global
— -
OpenAI’s new bidi voice mode seems crazy
By
–
OpenAI’s new upcoming „bidi“-voice mode sounds insane! pic.twitter.com/9UMzlCgEm9
— Chubby♨️ (@kimmonismus) 23 juin 2026OpenAI's upcoming new voice mode, the "bidi" mode, seems completely crazy!