8). Empowering MLLM with o1-like Reasoning and Reflection – proposes a new learning-to-reason method called CoMCTS that enables multimodal language models to develop step-by-step reasoning capabilities by leveraging collective knowledge from multiple models.
LLMS
-

ExploreToM Framework Reveals LLM Theory-of-Mind Limitations
By
–
6). Explore Theory-of-Mind – introduces ExploreToM, a framework that uses A* search to generate diverse, complex theory-of-mind scenarios that reveal significant limitations in current LLMs' social intelligence capabilities.
-
LLM Inference-Time Self-Improvement Survey Analysis
By
–
5). A Survey on LLM Inference-Time Self-Improvement – presents a survey that analyzes three categories of LLM inference-time self-improvement techniques – independent methods like enhanced decoding, context-aware approaches using external data, and model collaboration strategies.
-

ModernBERT: Efficient Encoder-Only Transformer for Classification Tasks
By
–
3). ModernBERT – a new encoder-only transformer model that achieves state-of-the-art performance on classification and retrieval tasks while being more efficient than previous encoders.
-
Large Concept Models: Beyond Token-Level Processing in LLMs
By
–
2). Large Concept Models – presents an approach that operates on sentence-level semantic representations called concepts, moving beyond token-level processing typical in current LLMs.https://t.co/GgonxsUMf0
— DAIR.AI (@dair_ai) 29 décembre 20242). Large Concept Models – presents an approach that operates on sentence-level semantic representations called concepts, moving beyond token-level processing typical in current LLMs.
-
Top ML Papers Week: DRT-o1, LearnLM, DeepSeek-V3
By
–
Here are the top ML Papers of the Week (Dec 16-22): – DRT-o1
– LearnLM
– DeepSeek-V3
– Large Concept Models
– Explore Theory-of-Mind
– Reinforcement Learning Overview Read on for more: -

DeepSeek-V3: 671B MoE Language Model with Efficient Parameter Activation
By
–
1). DeepSeek-V3 – a 671B-parameter MoE language model that activates 37B parameters per token, utilizing MLA and DeepSeekMoE architectures for efficient operation
-
Definition of AI Agents Involves Tool Use and Multistep Memory
By
–
@minimaxir i'd say agents are tool use + multistep (multistep means a memory that logs past steps ans errors properly)
-
Groq Reviewing New Models for Future Quarters
By
–
The team is actively reviewing a wide variety of models for addition in the coming quarters. If you would like to discuss any suggestions with the team please join our Discord and share your thoughts.
-
Anticipating Alpaca Moment for Thought Tree Datasets in 2025
By
–
100%! Particularly, I am hoping for an Alpaca moment for "thought tree" datasets in 2025.