Can AI finally grasp the unspoken intentions and emotions that make human social interactions tick? MetaMind helps large language models bridge that gap by breaking social reasoning into three collaborative stages: first guessing a user’s mental state, then refining those ideas
@jiqizhixin
-

LabOS: AI-Powered Co-Scientist Integrating XR for Research
By
–
The world's first Co-Scientist integrating AI and XR (Extended Reality)! Meet LabOS. It uses multimodal perception, self-evolving agents, and XR tools to see what researchers see, grasp experimental context, and assist in real time. From cancer immunotherapy target discovery
-

LoopTool: LLMs Learn Tool Use Through Adaptive Data Synthesis
By
–
What if LLMs could learn tool use faster by fixing their own training data gaps? SJTU & Xiaohongshu just built LoopTool for this. It closes the loop between data synthesis and model training, using the model’s own weaknesses to refine datasets adaptively. It probes what the
-

PH-Reg: Post-hoc Method Fixes Vision Transformer Artifacts
By
–
Ever wondered how to fix the weird artifact tokens that mess up Vision Transformers’ fine-grained tasks—without retraining those massive models from scratch? HKU, Zhejiang & NTU introduce PH-Reg, a post-hoc method that adds register tokens to pre-trained ViTs via
-

BraInCoRL: Transformer Models Visual Cortex Activity In-Context
By
–
What if we could model human visual cortex responses without expensive, time-consuming fMRI data for every new subject? This study introduces BraInCoRL, a transformer that uses in-context learning to predict voxelwise neural activity from just a few examples—no extra finetuning
-

ConsistEdit: Training-Free Text-Guided Image Video Editing
By
–
Ever wondered why text-guided image/video edits often lose source consistency or fail at fine-grained tweaks? This work on MM-DiT’s attention mechanisms unlocks a solution: ConsistEdit, a training-free method that balances strong editing strength with reliability. It uses
-

Video Models Zero-Shot Visual Reasoning Capabilities Study
By
–
Are current video models ready for zero-shot visual reasoning? A comprehensive empirical study on Veo-3 across 12 reasoning dimensions shows a mixed picture. These models handle short-horizon spatial coherence, local grounding, and simple dynamics, yet fail on long-horizon
-

Video-As-Prompt: Semantic Video Generation Without Retraining
By
–
What if video generation could follow any semantic instruction without retraining or task-specific hacks? Enter Video-As-Prompt (VAP). By treating a reference video as an in-context semantic prompt and steering a frozen DiT with a plug-and-play MoT expert plus temporally
-
DeepSeek-Math-V2: Advanced Mathematical Reasoning Model
By
–
deepseek-ai/DeepSeek-Math-V2 · Hugging Face
-

DeepSeek-Math-V2 Achieves Self-Verifiable Mathematical Reasoning
By
–
DeepSeek just released DeepSeek-Math-V2! DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning It shows LLMs can now self-verify proofs, not just output solutions. DeepSeekMath-V2 achieves gold-level IMO 2025, CMO 2024, and 118/120 Putnam 2024, pointing to a future of
