How can we make AI action generation both faster and more expressive without compromise? Researchers from Tsinghua University, Berkeley AI Research (BAIR), and The University of Hong Kong unveil their new Mean Velocity Policy (MVP). This innovative method models the "mean
@jiqizhixin
-

OPUS: Intelligent Data Selection for LLM Pre-training
By
–
Is there a smarter way to pick data for training Large Language Models? Researchers from multiple institutions, led by Shaobo Wang, introduce OPUS. This novel method dynamically and intelligently selects the most impactful data for LLM pre-training in every single training
-

HyperOffload: Compiler Framework Optimizes LLM Memory Management
By
–
Tired of LLMs maxing out memory even on powerful supernode architectures? Shanghai Jiao Tong University and Huawei Technologies Co., Ltd. introduce HyperOffload! This new compiler-assisted framework intelligently plans data movement for large language models. By treating
-

SWE-Vision: Teaching AI to Code Visual Intelligence
By
–
Can we unlock unprecedented visual intelligence in AI by teaching it to code what it sees? Researchers from UniPat AI and Michigan State University introduce SWE-Vision. This innovative agent allows vision-language models to write and execute Python code, using libraries like
-

MIT Reveals Simpler Method for Adapting Large AI Models
By
–
What if adapting large AI models for specific tasks was far simpler than we thought? Yulu Gan and Phillip Isola at MIT CSAIL reveal a surprising truth. They found that big pretrained models are already 'dense' with specialized experts. Their RandOpt method skips complex
-

V²Drop: Fast Large Vision-Language Models Without Accuracy Loss
By
–
How can we make Large Vision-Language Models (LVLMs) lightning fast without sacrificing accuracy? Researchers from Sichuan University, EPIC Lab @ Shanghai Jiao Tong University, and Zhejiang University have found a breakthrough! They introduce V²Drop, a new technique that
-

LoGeR: Building Consistent 3D Models from Video
By
–
How do you build a perfectly consistent 3D model from hours of video? Researchers from Google DeepMind and UC Berkeley unveil LoGeR. This new architecture processes video in chunks, using a clever hybrid memory system. It combines a global memory for overall scene consistency
-

NaLaFormer: Faster AI Models with Less Memory
By
–
What if your AI models could process massive data faster, with far less memory, and achieve state-of-the-art accuracy? Researchers from Harbin Institute of Technology, Pengcheng Laboratory, and UQMM Lab introduce NaLaFormer. Their NaLaFormer method uses a clever
-

MME-Emotion: Benchmarking Emotional Intelligence in Advanced AI Models
By
–
How truly emotionally intelligent are our most advanced AI models? A massive collaboration led by Fan Zhang et al. from The Chinese University of Hong Kong, Tongyi Lab, SZTU, and Tencent, introduces MME-Emotion. This groundbreaking benchmark, the largest of its kind, uses over
-
Google TurboQuant Team Faces Academic Misconduct Allegations
By
–
The Google TurboQuant team may be suspected of academic misconduct. This is concerning and warrants further attention.
