9. Visual Structures Help Visual Reasoning This study shows that adding simple spatial structures (like horizontal lines) to images significantly boosts GPT-4o’s visual reasoning by improving feature binding.
@dair_ai
-

Deep Research Agents: Comprehensive Survey of LLM-Powered Systems
By
–
6. Deep Research Agents Provides the most comprehensive survey to date of Deep Research (DR) agents, LLM-powered systems built for autonomous, multi-step informational research.
-

Evaluating LLM-Based Agents: Key Metrics and Methods
By
–
7. Survey on Evaluation of LLM-based Agents Overview of how to evaluate LLM-based agents, which differ significantly from traditional LLMs by maintaining memory, planning over multiple steps, using tools, and interacting with dynamic environments.
-

Comprehensive Threat Model for LLM-Powered AI Agent Ecosystems
By
–
5. Threats in LLM-Powered AI Agents Workflows This work presents the first comprehensive, end-to-end threat model for LLM-powered agent ecosystems.
-
Top AI Papers of the Week: Agents and LLM Research
By
–
Top AI Papers of The Week (June 30 – July 6): – xLSTMAD
– AI4Research
– Deep Research Agents
– SLMs are the Future of Agentic AI
– Chain-of-Thought Is Not Explainability
– Survey on Evaluation of LLM-based Agents Read on for more: -

Small Language Models: Superior Future of Agentic AI
By
–
1. Small Language Models are the Future of Agentic AI This position paper argues that small language models (SLMs), defined as those runnable on consumer-grade hardware, are not only sufficient but superior for many agentic AI applications.
-
PEVA: Egocentric Video Prediction from 3D Body Motion
By
–
10. Whole-Body Conditioned Egocentric Video Prediction
— DAIR.AI (@dair_ai) 29 juin 2025
This paper introduces PEVA, a conditional diffusion transformer that predicts egocentric video conditioned on 3D human body motion.https://t.co/CnDXcQHUqy10. Whole-Body Conditioned Egocentric Prediction This paper introduces PEVA, a conditional diffusion transformer that predicts egocentric video conditioned on 3D human body motion.
-

Security in LLM-Driven Agent Communication Protocols
By
–
8. AI Agent Communication Protocols Presents the first comprehensive survey on security in LLM-driven agent communication, categorizing it into three stages: user-agent interaction, agent-agent communication, and agent-environment communication.
-
Diffusion Steering via Reinforcement Learning for Policy Adaptation
By
–
9. Diffusion Steering via RL
— DAIR.AI (@dair_ai) 29 juin 2025
This paper introduces Diffusion Steering via Reinforcement Learning (DSRL), a method for adapting pretrained diffusion policies by learning in their latent-noise space instead of finetuning model weights.https://t.co/qDZKfUIIGv9. Diffusion Steering via RL This paper introduces Diffusion Steering via Reinforcement Learning (DSRL), a method for adapting pretrained diffusion policies by learning in their latent-noise space instead of finetuning model weights.
-

Anthropic Studies Emotional Support Seeking in Claude Conversations
By
–
7. Claude for Affective Use Anthropic presents the first large-scale study of how users seek emotional support from its Claude assistant, analyzing over 4.5 million conversations.
