Evaluating Multi-Modal Retrieval-Augmented Generation https://
bit.ly/49Xe6dB
#AI #MachineLearning #DeepLearning #LLMs #DataScience
MULTIMODAL AI
-

Evaluating Multi-Modal Retrieval-Augmented Generation Systems
By
–
-
Training ControlNet Models with Upscaled Video Game Screenshots
By
–
Could you use all these upscaled video game screenshots to train a controlnet?
-
AI Vision Limitations with Dynamic Movement and Scenes
By
–
I agree. I also wonder how it will deal with movements like sitting down, lying down, running – I imagine it only works well with static scenery.
-
Error Correction Codes in Language versus Video Models
By
–
Good point. But, is the redundancy in language the same as in video? Do both contain the same type of error-correction information? We can design for better error correction codes, so I’m not sure they’re the same
-
Auto-regressive Prediction Limits: Video vs Language Models
By
–
Thanks, Chris. Auto-regressive prediction has failed for video in the past but not obviously for language. @ylecun has emphasised this drift, but I wonder how much (1) redundancy and (2) discreteness can provide self-correction.
-
Winter Memories: AI Art Video Experiment with Midjourney
By
–
— Nathan Lands (@NathanLands) 2 décembre 2023
Winter Memories // a short video experiment made with
@midjourney + @runwayml #midjourney #AIart #aiartcommunity -

ART-V: Auto-Regressive Text-to-Video Generation Diffusion Model
By
–
ART⋅V: Auto-Regressive Text-to-Video Generation with Diffusion Models Weng et al.: https://
arxiv.org/abs/2311.18834 #ArtificialIntelligence #DeepLearning #MachineLearning -
Discreteness Redundancy Self-Correction Images Language Genes
By
–
And what role does discreteness play here? Images are also redundant but maybe not in the same way as English. So I’m not sure which accommodates better self-correction. I’m thinking of genes too. Clear answers appreciated.
-
SVD Impresses with Accurate 3D Image Generation
By
–
Holidays are coming.
— fofr (@fofrAI) 1 décembre 2023
So impressed how accurately SVD has figured out the 3D in this image. The trees, the background buildings, the windows in the gingerbread truck. pic.twitter.com/hRjCE59MaTHolidays are coming. So impressed how accurately SVD has figured out the 3D in this image. The trees, the background buildings, the windows in the gingerbread truck.
-

SeamlessExpressive: AI Speech-to-Speech Translation Preserving Voice
By
–
SeamlessExpressive: speech-to-speech translation that preserves the voice, the tone, and the expression. https://t.co/QJ8mTla8mE
— Yann LeCun (@ylecun) 1 décembre 2023SeamlessExpressive: speech-to-speech translation that preserves the voice, the tone, and the expression.