8/ Lumiere – a text-to-video space-time diffusion model for synthesizing videos with realistic and coherent motion; introduces a Space-Time U-Net architecture to generate the entire temporal duration of a video at once via a single pass. https://
x.com/GoogleAI/statu
s/1751003814931689487?s=20
…
MULTIMODAL AI
-
Lumiere: Space-Time Diffusion Model for Text-to-Video Synthesis
By
–
-

Resource-Efficient LLMs and Multimodal Models: Architecture and Implementation
By
–
6/ Resource-efficient LLMs & Multimodal Models – provides a comprehensive analysis and insights into ML efficiency research, including architectures, algorithms, and practical system designs and implementations.
-

Red Teaming Visual Language Models for Safety Alignment
By
–
7/ Red Teaming Visual Language Models – finds that 10 prominent open-sourced VLMs struggle with red teaming and have up to 31% performance gap with GPT-4V; applies red teaming alignment to LLaVA-v1.5 with SFT to improve performance by 10%.
-
Diffuse to Choose: Diffusion-Based Image Inpainting Model
By
–
4/ Diffuse to Choose – a diffusion-based image-conditioned inpainting model to balance fast inference with high fidelity while enabling accurate semantic manipulations in a given scene content.https://t.co/KUJjxVfCHc
— DAIR.AI (@dair_ai) 28 janvier 20244/ Diffuse to Choose – a diffusion-based image-conditioned inpainting model to balance fast inference with high fidelity while enabling accurate semantic manipulations in a given scene content.
-
Depth Anything: Monocular Depth Estimation from Unlabeled Data
By
–
1/ Depth Anything – a monocular depth estimation solution that can deal with any images under any circumstance; proposes effective strategies to leverage the power of the large-scale unlabeled data (~62M) which helps to reduce generalization error.https://t.co/cOXWWextRV
— DAIR.AI (@dair_ai) 28 janvier 20241/ Depth Anything – a monocular depth estimation solution that can deal with any images under any circumstance; proposes effective strategies to leverage the power of the large-scale unlabeled data (~62M) which helps to reduce generalization error.
-
Top Machine Learning Papers Week January 22-28
By
–
The Top ML Papers of the Week (Jan 22 – Jan 28): – WARM
– Medusa
– AgentBoard
– MambaByte
– Knowledge Fusion of LLMs
– Resource-efficient LLMs & Multimodal Models
… -
LLMs Face Text Data Scarcity, Video Data Offers New Opportunities
By
–
LLMs may be running out of text data: they've already scraped most of the web.
— Erik Brynjolfsson (@erikbryn) 28 janvier 2024
But beyond that, there are many other types of data, notably video. But that will require new approaches.@DaphneKoller and @ylecun explain it well in this short clip.https://t.co/spaJV6Io3YLLMs may be running out of text data: they've already scraped most of the web. But beyond that, there are many other types of data, notably video. But that will require new approaches. @DaphneKoller and @ylecun explain it well in this short clip.
-
Video datasets and emerging approaches as next frontier
By
–
Yep. And there are other approaches, using the much, much larger data sets of video (and other types of data) that could be the next frontier.
-
Optical Illusion Dress AI Generation Freaks Out Users Online
By
–
The black blue dress never worked on me.
— Andreas Klinger 🦾 (@andreasklinger) 27 janvier 2024
This one freaks me out though 🤯
I had to manually rewind to make sure it’s not two videos in one. 💀pic.twitter.com/6DNJdrgyTOThe black blue dress never worked on me. This one freaks me out though I had to manually rewind to make sure it’s not two videos in one.
-
ChatGPT adds granular controls for DALL-E image generation
By
–
ChatGPT integrates enhanced DALL-E controls for granular image customization WDYT?
Useful: Like Not useful: Reply