Meta-Personalizing Vision-Language Models to Find Named Instances in paper page: https://
huggingface.co/papers/2306.10
169
… Large-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they
MULTIMODAL AI
-

Meta-Personalizing Vision-Language Models for Named Instance Video Search
By
–
-
Robots learn by watching videos
By
–
Robots have developed the skill to learn by watching videos:
— AI Breakfast (@AiBreakfast) 20 juin 2023
The Visual-Robotics Bridge (VRB) model can teach a robot a new task in roughly 25 minutes using only reference video – without needing human interaction or a simulated environment. pic.twitter.com/iRPua6iTZaRobots have developed the skill to learn by watching videos: The Visual-Robotics Bridge (VRB) model can teach a robot a new task in roughly 25 minutes using only reference video – without needing human interaction or a simulated environment.
-
SoTA Audio Encoding for TTS/TTA Language Modeling Pipeline
By
–
Yep! It’s SoTA/ near-SoTA atm! This serves as the key piece in a TTS/ TTA pipeline. As it allows us to encode an audio into discrete representations and then perform language modelling on it!
-
Encodec: The Key Technology Behind MusicGen Audio Processing
By
–
Not quite, Encodec is part of the pipeline for MusicGen. Encodec helps with converting the audio into discrete codebook representation and back! Not as glamorous but it is the key piece behind making MusicGen as effective as it is
-

HierVL: Hierarchical Video-Language Embedding for Temporal Associations
By
–
HierVL is a novel hierarchical video-language embedding that simultaneously accounts for both long-term and short-term associations. Paper https://
bit.ly/3qSmEk6 7/7 -
Learning Video Representations from Large Language Models
By
–
📺 Learning Video Representations from Large Language Models
— AI at Meta (@AIatMeta) 20 juin 2023
This work repurposed pre-trained LLMs to be conditioned on visual input, and finetune them to create automatic video narrators.
Paper ➡️ https://t.co/OsZt3AnAEM
3/7 pic.twitter.com/DLunz5bcgPLearning Representations from Large Language Models This work repurposed pre-trained LLMs to be conditioned on visual input, and finetune them to create automatic video narrators. Paper https://
bit.ly/3NDCWpR 3/7 -
Transformers Integration Enables Large-Scale Text-to-Speech Music Models
By
–
Okay, but why is it a big deal!? Transformers integration allows us to use any LM and dataset from the ecosystem seamlessly to train Text-to-Speech and Text-to-Music models at scale! More exciting announcements on this front soon! https://
github.com/Vaibhavs10/not
ebooks/blob/main/use_encodec_w_transformers.ipynb
… -
EnCodec Model Now Available in Transformers Library
By
–
Want to train your own Bark/MusicGen-like TTS/TTA models? 👀
— Vaibhav (VB) Srivastav (@reach_vb) 20 juin 2023
The SoTA Encodec model by @MetaAI has now landed in 🤗Transformers!
It supports compression up to 1.5KHz and produces discrete audio representations. ⚡️
Model: https://t.co/Hq8rDHBfjw
Colab: https://t.co/MaWVEAMCXs pic.twitter.com/tDpPAdlYHUWant to train your own Bark/MusicGen-like TTS/TTA models? The SoTA Encodec model by @MetaAI has now landed in Transformers! It supports compression up to 1.5KHz and produces discrete audio representations. Model: https://
huggingface.co/docs/transform
ers/main/en/model_doc/encodec#overview
…
Colab: https://
github.com/Vaibhavs10/not
ebooks/blob/main/use_encodec_w_transformers.ipynb
… -

Infinigen: Procedural Generation for Photorealistic 3D Worlds
By
–
Infinite Photorealistic Worlds using Procedural Generation
— AK (@_akhaliq) 20 juin 2023
paper page: https://t.co/HiSl313uEN
introduce Infinigen, a procedural generator of photorealistic 3D scenes of the natural world. Infinigen is entirely procedural: every asset, from shape to texture, is generated from… pic.twitter.com/K1iGHwRahtInfinite Photorealistic Worlds using Procedural Generation paper page: https://
huggingface.co/papers/2306.09
310
… introduce Infinigen, a procedural generator of photorealistic 3D scenes of the natural world. Infinigen is entirely procedural: every asset, from shape to texture, is generated from -
TryOnGAN Image Inpainting: Exploring Generative AI Applications
By
–
Do you want to try TryOnGAN and do a sort of image inpainting? That would be super valuable 🙂