6/ More Agents is All You Need – presents a study on the scaling property of raw agents instantiated by LLMs; finds that performance scales when increasing agents by simply using a sampling-and-voting method.
@dair_ai
-
ALOHA 2: Advanced Bimanual Teleoperation System Improvements
By
–
5/ ALOHA 2 – a low-cost system for bimanual teleoperation that improves the performance, user-friendliness, and durability of ALOHA; efforts include hardware improvements such as grippers and gravity compensation with a higher quality simulation model.https://t.co/Xr7TxmqLxX
— DAIR.AI (@dair_ai) 11 février 20245/ ALOHA 2 – a low-cost system for bimanual teleoperation that improves the performance, user-friendliness, and durability of ALOHA; efforts include hardware improvements such as grippers and gravity compensation with a higher quality simulation model.
-

Indirect Reasoning Strengthens LLM Logical Reasoning Power
By
–
4/ Indirect Reasoning with LLMs – proposes an indirect reasoning method to strengthen the reasoning power of LLMs; it employs the logic of contrapositives and contradictions to tackle IR tasks such as factual reasoning and mathematic proof.
-

Phase Transition in Dot-Product Attention: Positional vs Semantic Learning
By
–
3/ A Phase Transition between Positional and Semantic Learning in a Solvable Model of Dot-Product Attention – studies the theoretical understanding of learning with attention layers by exploring the interplay between positional and semantic attention.
-

AnyTool: LLM Agent Framework for 16K API Integration
By
–
2/ AnyTool – an LLM-based agent that can utilize 16K APIs using a simple framework consisting of 1) a hierarchical API-retriever to identify relevant API candidates to a query, 2) a solver to resolve user queries, and 3) a self-reflection mechanism.
-

AgentBoard: Benchmark Framework for LLM Agent Evaluation
By
–
10/ AgentBoard – a benchmark with an open-source evaluation framework to perform analytical evaluation of LLM agents; assesses the capabilities and limitations of LLM agents and demystifies agent behaviors which leads to building stronger LLM agents.
-
Lumiere: Space-Time Diffusion Model for Text-to-Video Synthesis
By
–
8/ Lumiere – a text-to-video space-time diffusion model for synthesizing videos with realistic and coherent motion; introduces a Space-Time U-Net architecture to generate the entire temporal duration of a video at once via a single pass. https://
x.com/GoogleAI/statu
s/1751003814931689487?s=20
… -
Medusa Framework Achieves 2.2x LLM Inference Speedup
By
–
9/ Medusa – a framework for LLM inference acceleration using multiple decoding heads; substantially reduces the number of decoding steps and achieves over 2.2x speedup without compromising generation quality.
-

Resource-Efficient LLMs and Multimodal Models: Architecture and Implementation
By
–
6/ Resource-efficient LLMs & Multimodal Models – provides a comprehensive analysis and insights into ML efficiency research, including architectures, algorithms, and practical system designs and implementations.
-

Red Teaming Visual Language Models for Safety Alignment
By
–
7/ Red Teaming Visual Language Models – finds that 10 prominent open-sourced VLMs struggle with red teaming and have up to 31% performance gap with GPT-4V; applies red teaming alignment to LLaVA-v1.5 with SFT to improve performance by 10%.
