From more and more advanced reasoning models to reconstructing the world in 3D—AI is reshaping reality. What a way to cap the year! – Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
– Align Anything: Training All-Modality
@askalphaxiv
-

Advanced Reasoning Models and Multimodal AI Reshape Reality
By
–
-

Open-Source Advantage in Large Language Models Debate
By
–
The Open-Source Advantage in Large Language Models (LLMs) Large Language Models (LLMs) are revolutionizing natural language processing, but the debate between closed-source and open-source models raises important questions about transparency, accessibility, and ethics.
-

LLM Agents Evolve Cooperation Through Indirect Reciprocity
By
–
Cultural Evolution of Cooperation among LLM Agents A study examining how large language model (LLM) agents evolve cooperation and social norms over generations, specifically focusing on indirect reciprocity. Problem: Limited understanding of how multiple LLM agents interact
-

Grokking Complexity Dynamics: Memorization to Generalization
By
–
The Complexity Dynamics of Grokking A study on how neural network complexity dynamics explain the grokking phenomenon, where models transition from memorization to generalization long after overfitting. Problem: Grokking challenges our understanding of generalization in
-

Adaptive Computation Modules: Efficient Token-Level Conditional Inference
By
–
Adaptive Computation Modules: Granular Conditional Computation for Efficient Inference A neural network module that dynamically adapts computational load per token, reducing inference costs without sacrificing accuracy. Problem: Transformer models are computationally
-

AniDoc: AI-Powered 2D Animation Colorization and In-Betweening
By
–
AniDoc: Animation Creation Made Easier A tool leveraging generative AI to automate 2D animation colorization and in-betweening for more efficient animation production. Problem: The 2D animation process, especially line art colorization and in-betweening, is
-

LlamaFusion Extends LLMs with Multimodal Capabilities
By
–
LlamaFusion Extends pretrained text-only LLMs with multimodal capabilities while preserving language performance. Problem:
Training multimodal models from scratch is costly and risks degrading pretrained language abilities. Method:
Adds image-specific modules to pretrained -

MetaMorph: Multimodal LLM Extension via Instruction Tuning
By
–
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning A method to extend LLMs for unified visual and textual generation, leveraging instruction tuning for efficient multimodal adaptation. Problem: Unified models for visual understanding and generation
-

Best-of-N Jailbreaking Method Bypasses Major AI Models
By
–
Best-of-N Jailbreaking Introducing a black-box method that jailbreaks AI models across text, vision, and audio by applying simple prompt augmentations. BoN achieves high success rates, such as 89% on GPT-4o and 78% on Claude 3.5 Sonnet, with 10,000 augmented prompts, and also
-

LMAgent: Multimodal Large-Scale Multi-User Simulation Framework
By
–
LMAgent: A Large-Scale Multimodal Agents Society for Multi-user Simulation A scalable framework for simulating dynamic, multimodal multi-user behavior using large-scale multimodal LLMs. Problem: Existing LLM-based multi-agent systems simulate only text-based interactions