9). Adapting while Learning – proposes a two-part fine-tuning approach that first helps LLMs learn from tool-generated solutions and then trains them to determine when to solve problems directly versus when to use tools; testing on math, climate science, and epidemiology
@dair_ai
-
Personalization Framework for Large Language Models
By
–
10). Personalization of LLMs – presents a comprehensive framework for understanding personalized LLMs; introduces taxonomies for different aspects of personalization and unifying existing research across personalized text generation and downstream applications.
-

LLM Numerical Understanding: Fine-tuning Improves Processing Ability
By
–
7). Number Understanding of LLMs – provides a comprehensive analysis of the numerical understanding and processing ability (NUPA) of LLMs; finds that naive finetuning can improve NUPA a lot on many but not all tasks…
-
WebRL Self-Evolving Framework Boosts Open LLM Web Agents
By
–
8). WebRL – proposes a self-evolving online curriculum RL framework to bridge the gap between open and proprietary LLM-based web agents; it improves the success rate of Llama-3.1-8B from 4.8% to 42.4%, and from 6.1% to 43% for GLM4-9B; the open models significantly surpass the
-
Multi-Expert Prompting: Aggregating LLM Responses for Better Outputs
By
–
6). Multi-expert Prompting with LLMs – improves LLM responses by simulating multiple experts and aggregating their responses; it guides an LLM to fulfill input instructions by simulating multiple experts and selecting the best response among individual and aggregated views.
-

Vision-Language Agents Vulnerable to Pop-up Adversarial Attacks
By
–
5). Attacking Vision-Language Agents via Pop-ups – shows that integrating adversarial pop-ups into existing agent testing environments leads to an attack success rate of 86%; this decreases the agents' task success rate by 47%.
-

Mixtures of In-Context Learners: Expert Weighting for Token Prediction
By
–
4). Mixtures of In-Context Learners – uses subsets of demonstrations to train experts via in-context learning; given a training set, a trainable weighting function is used to combine the experts' next-token predictions…
-
OpenAI o1 Model Reasoning Patterns and Performance Analysis
By
–
10). Reasoning Patterns of OpenAI’s o1 Model – when compared with other test-time compute methods, o1 achieved the best performance across most datasets; the authors observe that the most commonly used reasoning patterns in o1 are divide and conquer and self-refinement.
-
SynthID-Text: Scalable Watermarking Scheme for LLMs
By
–
9). Scalable Watermarking for LLMs – proposes SynthID-Text, a text-watermarking scheme that can preserve text quality in LLMs, enable high detection accuracy, and minimize latency overhead…https://t.co/isvnHcP610
— DAIR.AI (@dair_ai) 27 octobre 20249). Scalable Watermarking for LLMs – proposes SynthID-Text, a text-watermarking scheme that can preserve text quality in LLMs, enable high detection accuracy, and minimize latency overhead…
-

LLMs Reflect Creator Ideology Across Languages
By
–
8). LLMs Reflect the Ideology of their Creators – finds that LLMs exhibit a diverse ideological stance which reflects the worldview of its creators; finds consistent normative differences between how the same LLM responds in Chinese compared to English.