7). Agent S – a new open agentic framework that enables autonomous interaction with computers through a GUI; Agent S tackles challenges such as acquiring knowledge, planning over long-task horizons, and handling dynamic interfaces.
@dair_ai
-
Janus: Unified Autoregressive Framework for Multimodal Understanding
By
–
5). Janus – proposes a unified autoregressive framework for multimodal understanding and generation; it decouples visual encoding into independent pathways and leverages a single transformer architecture to improve flexibility and performance on both visual understanding and
-

Inference Scaling Laws for Long-Context RAG Systems
By
–
6). Inference Scaling for Long-Context RAG – uses two strategies to investigate scaling laws for RAG: in-context learning and iterative prompting; RAG performance improves with the expansion of the effective context length under optimal configurations.
-
LLM Introspection Enables Self-Knowledge and System Interpretability
By
–
4). Introspection in LLMs – reports that LLMs can acquire knowledge through introspection that cannot be inferred from their training data; suggests that LLMs contain privileged information about themselves that can potentially lead to more interpretable and controllable systems.
-
First-Person Fairness Biases in ChatGPT Analysis
By
–
3). First-Person Fairness in Chatbots – studies first-person fairness which involves fairness towards users interacting with ChatGPT; specifically, it measures the biases, if any, towards the users’ names…
-
Designing Priors for Few-Shot Image Synthesis with GANs
By
–
10). Designing Priors for Better Few-Shot Image Synthesis – training generative models like GAN with limited data is difficult; current Implicit Maximum Likelihood Estimation approaches (IMLE) have an inadequate correspondence between latent code selected for training and those
-
Comprehensive Evaluation of OpenAI’s o1-preview LLM Model
By
–
9). Evaluation of o1 – provides a comprehensive evaluation of OpenAI's o1-preview LLM; shows strong performance across many tasks such as competitive programming, generating coherent and accurate radiology reports, high school-level mathematical reasoning tasks, chip design
-
LLM Reasoning Gaps in Grade-School Math Problem Solving
By
–
8). Not All LLM Reasoners Are Created Equal – investigates in depth the grade-school math problem-solving capabilities of LLMs; reports that LLMs show a significant gap in reasoning; finds that LLMs display a huge performance difference when solving compositional pairs and
-

o1-preview Analysis: Advanced Reasoning Models Show Similar LLM Trends
By
–
6). An Analysis of o1-preview – reports that large reasoning models like o1-preview, while improving on more difficult tasks, display similar qualitative trends as previous LLMs…
-

FRAMES: Framework for Evaluating LLM Factuality and Reasoning
By
–
7). FRAMES – a unified framework to evaluate an LLM’s ability to provide factual responses, assess retrieval capabilities, and the reasoning required to generate final responses…
