7). Granite 3.0 – presents lightweight foundation models ranging from 400 million to 8B parameters; supports coding, RAG, reasoning, and function calling, focusing on enterprise use cases, including on-premise and on-device settings.
@dair_ai
-
Feature Steering in LLMs: Evaluating Social Bias Control
By
–
6). Evaluation Feature Steering in LLMs – evaluates featuring steering in LLMs using an experiment that artificially dials up and down various features to analyze changes in model outputs; it focused on 29 features related to social biases and study if feature steering can help
-

LongRAG Enhances RAG Understanding Long-Context Knowledge
By
–
5). LongRAG – enhances RAG's understanding of long-context knowledge which includes global information and factual details; consists of a hybrid retriever, an LLM-augmented information extractor, a CoT-guided filter, and an LLM-augmented generator…
-
Data Synthesis and Augmentation Techniques for Large Language Models
By
–
4). A Survey on Data Synthesis and Augmentation for LLMs – provides a comprehensive summary of data generation techniques in the lifecycle of LLMs; includes discussions on data preparation, pre-training, fine-tuning, instruction-tuning, preference alignment, and applications.
-
Aya Expanse: Multilingual Foundation Models 8B and 32B Parameters
By
–
2). Aya Expanse – a family of open-weight foundation models for multilingual capabilities; releases an 8B and 32B parameter model, including one of the largest multilingual dataset collections to date, with 513 million examples… https://t.co/LWH49PJKZ0
— DAIR.AI (@dair_ai) 27 octobre 20242). Aya Expanse – a family of open-weight foundation models for multilingual capabilities; releases an 8B and 32B parameter model, including one of the largest multilingual dataset collections to date, with 513 million examples…
-

Agentic Information Retrieval: LLM Agent Applications
By
–
1). Agentic Information Retrieval – provides an introduction to agentic information retrieval, which is shaped by the capabilities of LLM agents; discusses different types of cutting-edge applications of agentic information retrieval and challenges.
-
Top ML Papers of the Week: LongRAG, Granite 3.0, and Beyond
By
–
The Top ML Papers of the Week (Oct 21 – 27): – LongRAG
– Granite 3.0
– Data Synthesis Overview
– Agentic Information Retrieval
– Scalable Watermarking for LLMs
– A Theoretical Understanding of CoT Read on for more: -
CoTracker3: Semi-Supervised Point Tracking Model with Pseudo-Labels
By
–
10). CoTracker3 – proposes a new point tracking model and a new semi-supervised training recipe; enables usage of real videos without annotations during training by generating pseudo-labels using off-the-shelf teachers.https://t.co/vCbJzcXSPj
— DAIR.AI (@dair_ai) 20 octobre 202410). CoTracker3 – proposes a new point tracking model and a new semi-supervised training recipe; enables usage of real videos without annotations during training by generating pseudo-labels using off-the-shelf teachers.
-
OpenAI o1 Models Planning Abilities Self-Evaluation Constraints
By
–
9). On the Planning Abilities of OpenAI’s o1 Models – reports that o1-preview is particularly strong in self-evaluation and constraint-following; also mentions that these o1 models demonstrate bottlenecks in decision-making and memory management, which are more pronounced in
-

Model Kinship Strategy for Improved LLM Merging
By
–
8). Model Kinship for Merging LLMs – proposes model kinship to measure the degree of similarity between LLMs; model kinship is used to build a model merging strategy (Top-k Greedy Merging with Model Kinship) which yields better performance.
