Top ML Papers of the Week (June 5-11): – AlphaDev
– MusicGen
– Fine-Grained RLHF
– Humor in ChatGPT
– Concept Scrubbing in LLM
– Augmenting LLMs with Databases
…
MULTIMODAL AI
-
Top ML Papers Week: AlphaDev, MusicGen, RLHF Advances
By
–
-
Dense Motion Estimation Method Tracks Pixels Across Full Videos
By
–
1/ Tracking Everything Everywhere All at Once – propose a test-time optimization method for estimating dense and long-range motion; enables accurate, full-length motion estimation of every pixel in a video.https://t.co/7O3Z0Em7wE
— DAIR.AI (@dair_ai) 11 juin 20231/ Tracking Everything Everywhere All at Once – propose a test-time optimization method for estimating dense and long-range motion; enables accurate, full-length motion estimation of every pixel in a video.
-
MusicGen, BAAI LLM, LIMA Dataset, Llama Falcon Variants
By
–
MusicGen, BAAI’s new LLM, LIMA dataset launch, a billion new variants of Llama/ Falcon & much more.
— Vaibhav (VB) Srivastav (@reach_vb) 11 juin 2023
You were missed hehe! 🤗 pic.twitter.com/rg2mF0hk7UMusicGen, BAAI’s new LLM, LIMA dataset launch, a billion new variants of Llama/ Falcon & much more. You were missed hehe!
-

Keanu Reeves as Superhero in Taiwanese Night Market via Midjourney
By
–
I asked for Keanu Reeves as a superhero in a Taiwanese night market buying snack. Greg Hildebrandt comic artist style. I guess AI knows Keanu doesn’t need a costume to be a superhero. #midjourney
-
NYC Rat Documentary Created with Text to Video AI Modelscope
By
–
NYC Rat Documentary, text to video AI, Modelscope pic.twitter.com/Zvtn44iVTy
— AK (@_akhaliq) 11 juin 2023NYC Rat Documentary, text to video AI, Modelscope
-

Word-As-Image for Semantic Typography in Machine Learning
By
–
Word-As-Image for Semantic Typography https://
bit.ly/3MI8bjk #MachineLearning #AI #DataScience #DeepLearning -
Real-Time Model Inference Streaming for Low-Latency AI Applications
By
–
Streaming is typically used when the data is arriving in real time, or when the latency requirements are very low. A typical example of this kind of model inference is the social media feeds, movie recsys, etc.
-

PandaGPT: Multimodal AI Model for Vision, Audio, Text Tasks
By
–
PandaGPT is a general-purpose instruction-following model that can both see and hear. PandaGPT can perform complex tasks such as detailed image description generation, writing stories inspired by videos, and answering questions about audios. https://
bit.ly/45TnAVl -
QR Codes Meet Stable Diffusion: AI Vision Evolution
By
–
I never imagined QR codes & Stable diffusion would collide. They seem to be from such different worlds! Funny story, 11 years ago, I tried to make a point that QR codes will be made useless thanks to computer vision: https://
youtube.com/watch?v=8iPgYV
-__mo
…. Guess I was wrong! -
Fine-tuning Vision Transformers shows promising results
By
–
I'd say so! Recently finetuned a couple of ViTs and it worked pretty well actually.