Scaling Speech-Text Pre-Training with Synthetic Interleaved Data A method for scaling speech language models (SpeechLMs) by using synthetic speech-text interleaved data, bypassing the need for parallel speech-text datasets. Problem: Limited unsupervised speech and parallel
@askalphaxiv
-

NVILA: Efficient Frontier Visual Language Models
By
–
NVILA: Efficient Frontier Visual Language Models A family of efficient and accurate visual language models that optimize processing high-resolution images and long videos. Problem: While VLMs have advanced in accuracy, their efficiency remains a significant challenge,
-

PaliGemma 2: Versatile Vision-Language Models for Transfer
By
–
PaliGemma 2: A Family of Versatile VLMs for Transfer An upgraded vision-language model family optimized for transfer tasks across various domains and resolutions. Problem: Previous VLMs lacked versatility in model size, resolution, and task-specific transfer capabilities,
-

AI Governance Blueprint: Maximizing Benefits While Managing Risks
By
–
Shaping AI's Impact on Billions of Lives A blueprint for leveraging AI to maximize societal benefits while minimizing risks through guided milestones and collaboration. Problem: AI’s transformative potential risks being guided solely by commercial interests, leading to
-

Motion Prompting Framework for Precise Video Generation Control
By
–
Motion Prompting for Generation A flexible motion control framework for generating videos with precise spatio-temporal dynamics through motion prompts. Problem: Existing video generation models lack nuanced control over dynamic actions and temporal compositions,
-

Florence-VL: Enhanced Multimodal Vision-Language Model
By
–
Florence-VL A multimodal large language model (MLLM) that integrates enriched visual representations from Florence-2 for improved vision-language alignment. Problem: Existing vision-language models (VL) are limited by less versatile visual representations and the need for
-

FLOAT: Expressive Talking Portrait Generation from Audio
By
–
FLOAT A method for generating expressive, temporally consistent talking portrait videos from a single image and audio. Problem: Audio-driven talking portrait generation faces challenges in creating temporally consistent motion and efficient sampling while maintaining
-

EfficientTAMs: Lightweight Video Object Segmentation Models
By
–
Efficient Track Anything (EfficientTAMs) Lightweight models for video object segmentation and tracking, optimized for speed and smaller model sizes. Problem: SAM 2's high computational complexity limits real-world applications, especially on mobile devices. Method:
-

Open-Sora Plan: Open-Source High-Resolution Video Generation
By
–
Open-Sora Plan: Open-Source Large Generation Model An open-source framework for generating high-resolution, long-duration videos from diverse user inputs. Problem: Existing video generation models struggle with high-resolution, long-duration video outputs due to high
-

Navigation World Models: Video Generation for Autonomous Agent Planning
By
–
Navigation World Models A controllable video generation model that predicts future visual observations for navigation tasks based on past observations and actions. Problem: Visual-motor agents struggle with planning flexible navigation trajectories, especially in dynamic or
