8). Empowering MLLM with o1-like Reasoning and Reflection – proposes a new learning-to-reason method called CoMCTS that enables multimodal language models to develop step-by-step reasoning capabilities by leveraging collective knowledge from multiple models.
@dair_ai
-
LearnLM: Adaptive Teaching Model with Pedagogical Instructions
By
–
7). LearnLM – a new LearnLM model that can follow pedagogical instructions, allowing it to adapt its teaching approach based on specified educational needs rather than defaulting to simply presenting information.
-

ExploreToM Framework Reveals LLM Theory-of-Mind Limitations
By
–
6). Explore Theory-of-Mind – introduces ExploreToM, a framework that uses A* search to generate diverse, complex theory-of-mind scenarios that reveal significant limitations in current LLMs' social intelligence capabilities.
-
LLM Inference-Time Self-Improvement Survey Analysis
By
–
5). A Survey on LLM Inference-Time Self-Improvement – presents a survey that analyzes three categories of LLM inference-time self-improvement techniques – independent methods like enhanced decoding, context-aware approaches using external data, and model collaboration strategies.
-
Foundation Models Automate Discovery of Artificial Life Simulations
By
–
4). Automating the Search for Artificial Life – presents a new approach that uses foundation models to automatically discover interesting artificial life simulations across multiple platforms like Boids, Lenia, and Game of Life.https://t.co/CQ0W161AE1
— DAIR.AI (@dair_ai) 29 décembre 20244). Automating the Search for Artificial Life – presents a new approach that uses foundation models to automatically discover interesting artificial life simulations across multiple platforms like Boids, Lenia, and Game of Life.
-

ModernBERT: Efficient Encoder-Only Transformer for Classification Tasks
By
–
3). ModernBERT – a new encoder-only transformer model that achieves state-of-the-art performance on classification and retrieval tasks while being more efficient than previous encoders.
-
Large Concept Models: Beyond Token-Level Processing in LLMs
By
–
2). Large Concept Models – presents an approach that operates on sentence-level semantic representations called concepts, moving beyond token-level processing typical in current LLMs.https://t.co/GgonxsUMf0
— DAIR.AI (@dair_ai) 29 décembre 20242). Large Concept Models – presents an approach that operates on sentence-level semantic representations called concepts, moving beyond token-level processing typical in current LLMs.
-
Top ML Papers Week: DRT-o1, LearnLM, DeepSeek-V3
By
–
Here are the top ML Papers of the Week (Dec 16-22): – DRT-o1
– LearnLM
– DeepSeek-V3
– Large Concept Models
– Explore Theory-of-Mind
– Reinforcement Learning Overview Read on for more: -

DeepSeek-V3: 671B MoE Language Model with Efficient Parameter Activation
By
–
1). DeepSeek-V3 – a 671B-parameter MoE language model that activates 37B parameters per token, utilizing MLA and DeepSeekMoE architectures for efficient operation
-
Mathematical Reasoning in Multimodal LLMs: Comprehensive Survey
By
–
9). A Survey of Mathematical Reasoning in the Era of Multimodal LLMs – presents a comprehensive survey analyzing mathematical reasoning capabilities in multimodal large language models (MLLMs), covering benchmarks, methodologies, and challenges across 200+ studies since 2021.