7/ TestGen-LLM – uses LLMs to improve existing human-written tests; reports that after an evaluation on Reels and Stories products for Instagram, 75% of TestGen-LLM's test cases were built correctly, 57% passed reliably, and 25% increased coverage.
@dair_ai
-

OS-Copilot Framework Builds Generalist Computer Agents
By
–
6/ OS-Copilot – a framework to build generalist computer agents that interface with key elements of an operating system like Linux or MacOS; it also proposes a self-improving embodied agent for automating general computer tasks.
-

Gemini 1.5: Compute-Efficient Multimodal Mixture-of-Experts Model
By
–
2/ Gemini 1.5 – a compute-efficient multimodal mixture-of-experts model with capabilities such as recalling and reasoning over long-form content; reasons over long documents potentially containing millions of tokens, including hours of video and audio.https://t.co/oZyOtMNBnR
— DAIR.AI (@dair_ai) 18 février 20242/ Gemini 1.5 – a compute-efficient multimodal mixture-of-experts model with capabilities such as recalling and reasoning over long-form content; reasons over long documents potentially containing millions of tokens, including hours of video and audio.
-

V-JEPA: Self-Supervised Vision Model Trained on Two Million Videos
By
–
3/ V-JEPA – vision models trained on a feature prediction objective using 2 million videos; relies on self-supervised learning and doesn’t use pretrained image encoders, text, negative examples, reconstruction, or other supervision sources.
-
Sora: Text-to-Video AI Model Generates Realistic Minute-Long Videos
By
–
1/ Sora – a text-to-video AI model that can create videos of up to a minute of realistic and imaginative scenes given text instructions; it can generate complex scenes with multiple characters, different motion types, and backgrounds, and understand how they relate to each other.
-
Top ML Papers Week: Sora, Gemini 1.5, Agents
By
–
The Top ML Papers of the Week (Feb 12 – Feb 18): – Sora
– Gemini 1.5
– OS-Copilot
– TestGen-LLM
– Large World Model
– LLM Agents can Hack
… -

LLM-Based Multi-Agent Systems: Applications, Benchmarks, and Challenges
By
–
10/ LLM-based Multi-Agents – discusses the essential aspects of LLM-based multi-agent systems; it includes a summary of recent applications for problem-solving and word simulation; summarizes datasets, benchmarks, challenges, and future opportunities.
-

DeepSeekMath Enhances Mathematical Reasoning with GRPO Optimization
By
–
8/ DeepSeekMath – continues pretraining a code base model with 120B math-related tokens; introduces GRPO (a variant to PPO) to enhance mathematical reasoning and reduce training resources via a memory usage optimization scheme.
-

Self-Discover: LLMs Select Reasoning Techniques for Tasks
By
–
7/ Self-Discovered Reasoning Structures – proposes a new framework, Self-Discover, that enables LLMs to select from multiple reasoning techniques (e.g., critical thinking and thinking step-by-step) to compose task-specific reasoning strategies.
-

LLMs for Table Processing: Methods, Benchmarks and Techniques
By
–
9/ LLMs for Table Processing – provides an overview of LLMs for table processing, including methods, benchmarks, prompting techniques, and much more.
