5. CoAct-1 Researchers from USC, Salesforce, and UW present CoAct-1, a multi-agent system that combines GUI interaction with direct code execution to improve efficiency and robustness in computer-using agents.
@dair_ai
-

Seed Diffusion: Discrete-State LLM for Code Generation
By
–
6. Seed Diffusion A discrete-state diffusion-based LLM optimized for code generation, achieving 2,146 tokens/sec on H20 GPUs while maintaining competitive benchmark performance.
-
Apple Launches Multilingual Foundation Models for On-Device AI
By
–
10. Apple Intelligence Foundation Language Models Apple introduces two multilingual, multimodal foundation models: a 3B-parameter on-device model optimized for Apple silicon and a scalable server model using a novel PT-MoE transformer architecture.
-
MCPEval: Open-Source LLM Agent Evaluation Framework
By
–
9. MCPEval MCPEval is an open-source framework that automates end-to-end evaluation of LLM agents using a standardized Model Context Protocol, eliminating manual benchmarking.
-

Deep Researcher: Test-Time Diffusion for Report Generation
By
–
8. Deep Researcher with Test-Time Diffusion Rather than relying on static inference strategies like CoT or best-of-n sampling, this work frames the report generation process as a diffusion process.
-

Inverse Scaling in Large Reasoning Models Test-Time Compute
By
–
6. Inverse Scaling in Test-Time Compute Presents a systematic study of inverse scaling in large reasoning models, where increasing the test-time compute (i.e., reasoning length) harms rather than helps model performance.
-
Compute-Optimal Strategies for Many-Shot In-Context Learning
By
–
7. Towards Compute-Optimal Many-Shot In-Context Learning Proposes practical strategies for reducing the cost of many-shot in-context learning while preserving or improving performance. https://
arxiv.org/abs/2507.16217 -

In-Context Learning Without Weight Updates in LLMs
By
–
5. Learning without Training This paper provides a theoretical and empirical explanation for how LLMs exhibit in-context learning, the ability to learn from examples in a prompt without weight updates.
-

Routine: Structured Planning Format for LLM Agent Tool-Calling
By
–
4. Structural Planning for LLM Agents This paper introduces Routine, a structured planning format designed to improve the stability and accuracy of LLM agents executing multi-step tool-calling tasks in enterprise settings.
-
LLMs in AIOps: Survey of 183 Papers and Methods
By
–
10. A Survey of AIOps This survey analyzes 183 papers to evaluate how LLMs are being used in AIOps, focusing on data sources, task evolution, applied methods, and evaluation practices.
