Aladdin: Zero-Shot Hallucination of Stylized 3D Assets from Abstract Scene Descriptions paper page: https://
huggingface.co/papers/2306.06
212
… What constitutes the "vibe" of a particular scene? What should one find in "a busy, dirty city street", "an idyllic countryside", or "a crime scene in an
MULTIMODAL AI
-

Aladdin: Zero-Shot 3D Asset Generation from Scene Descriptions
By
–
-
Meta Scales Speech Technology to 1,100+ Languages Globally
By
–
Our work on the Massively Multilingual Speech (MMS) project has scaled speech-to-text & text-to-speech to support 1,100+ languages + trained new language identification models that can identify 4,000+ languages!
— AI at Meta (@AIatMeta) 12 juin 2023
More info & access to pretrained models ⬇️Our work on the Massively Multilingual Speech (MMS) project has scaled speech-to-text & text-to-speech to support 1,100+ languages + trained new language identification models that can identify 4,000+ languages! More info & access to pretrained models
-
AI Generated Video Using Runway Gen-1, Elevenlabs and Reface
By
–
AI generated video made with Runway Gen-1, Elevenlabs & Reface by @Martin_Haerlinpic.twitter.com/SYUhtqLHk7
— AK (@_akhaliq) 12 juin 2023AI generated video made with Runway Gen-1, Elevenlabs & Reface by @Martin_Haerlin
-

Advanced Computer Vision Course at UCF Center for Research
By
–
Advanced Computer Vision – UCF Center for Research in Computer Vision, Spring 2023 This is an excellent course on advanced computer vision. Covers recent developments in computer vision, topics like image generation and vision-language learning. I like that the course go
-
PaLI-X Combines ViT-22B Vision and Multilingual UL2-32B
By
–
Check out PaLI-X that combines ViT-22B and multilingual UL2-32B!
-
Coronavirus as a Pokemon – Weird DALLE Reddit Thread
By
–
reddit thread: https://reddit.com/r/weirddalle/comments/147bnuq/coronavirus_as_a_pokemon/
-
Adapter v2 and Video-LLaMA Models Available
By
–
There's both, adapter v2 and e.g., Video-LLaMA https://
github.com/DAMO-NLP-SG/Vi
deo-LLaMA
… -
Open Source Models Now Support Multiple Media Types Beyond Images
By
–
Yeah, in the demo they showed a version that supported images. But still, afaik it was only images whereas there are now open source models that also support audio, video, etc.
-

Create AI Text-to-Video Clips in Seconds
By
–
How to Create an AI Text-to-Video Clip in Seconds – dubious seed data ! https://
buff.ly/3ClyNjW
#ai #artificialintelligence #MachineLearning #DeepLearning #GenerativeAI #video
