Marketing AI Show ep 49 was packed w/ news from @GoogleAI (Ads) @Microsoft (CoPilots) @DeepMind (risk warning systems) @MetaAI (multilingual speech project) @AnthropicAI ($$$), @Figure_robot ($$$) @worldnetwork (Sam Altman’s other company) @nvidia ($$$).
MULTIMODAL AI
-

Stability AI releases Pick-a-Pic dataset for image generation
By
–
Researchers at Stability AI helped collect and create Pick-a-Pic – an open dataset of user preferences of AI-generated images. This dataset will enable improvements in future text-to-image models! Learn more on our new research blog: https://
stability.ai/research/pick-
a-pic
… -
Humans in 4D: Reconstructing and Tracking Humans with Transformers
By
–
Humans in 4D: Reconstructing and Tracking Humans with Transformers
— AK (@_akhaliq) 1 juin 2023
present an approach to reconstruct humans and track them over time. At the core of our approach, we propose a fully "transformerized" version of a network for human mesh recovery. This network, HMR 2.0, advances… pic.twitter.com/46FkK7WHgFHumans in 4D: Reconstructing and Tracking Humans with Transformers present an approach to reconstruct humans and track them over time. At the core of our approach, we propose a fully "transformerized" version of a network for human mesh recovery. This network, HMR 2.0, advances
-
Markerless Motion Capture Revolutionizes Immersive Sports Experience
By
–
Dive into a new era of immersive sports with Guillaume Chican's groundbreaking tech! ⚽
— Pascal Bornet (@pascal_bornet) 1 juin 2023
An #innovation harnessing real-time markerless mocap technology to virtually recreate the action! Brought to life on Tilt Five #Holographic system
Credit: R. Emig#ar #tech #ai #vr pic.twitter.com/GJfHBfrnSfDive into a new era of immersive sports with Guillaume Chican's groundbreaking tech! An #innovation harnessing real-time markerless mocap technology to virtually recreate the action! Brought to life on Tilt Five #Holographic system Credit: R. Emig
#ar #tech #ai #vr -

Improving CLIP Training with Language Rewrites via LaCLIP
By
–
Improving CLIP Training with Language Rewrites introduce Language augmented CLIP (LaCLIP), a simple yet highly effective approach to enhance CLIP training through language rewrites. Leveraging the in-context learning capability of large language models, we rewrite the text
-

Concept Decomposition for Visual Exploration Using Vision-Language Models
By
–
Concept Decomposition for Visual Exploration and Inspiration
— AK (@_akhaliq) 31 mai 2023
propose a method to decompose a visual concept, represented as a set of images, into different visual aspects encoded in a hierarchical tree structure. We utilize large vision-language models and their rich latent… pic.twitter.com/J5OduSX7CGConcept Decomposition for Visual Exploration and Inspiration propose a method to decompose a visual concept, represented as a set of images, into different visual aspects encoded in a hierarchical tree structure. We utilize large vision-language models and their rich latent
-

VisorGPT: Learning Visual Prior via Generative Pre-Training
By
–
VisorGPT: Learning Visual Prior via Generative Pre-Training
— AK (@_akhaliq) 31 mai 2023
propose to learn Visual prior via Generative Pre-Training, dubbed VisorGPT. By discretizing visual locations of objects, e.g., bounding boxes, human pose, and instance masks, into sequences, our~can model visual prior… pic.twitter.com/xj84MvpE14VisorGPT: Learning Visual Prior via Generative Pre-Training propose to learn Visual prior via Generative Pre-Training, dubbed VisorGPT. By discretizing visual locations of objects, e.g., bounding boxes, human pose, and instance masks, into sequences, our~can model visual prior
-

LibriTTS-R: Restored Multi-Speaker Text-to-Speech Dataset
By
–
LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus paper introduces a new speech dataset called “LibriTTS-R'' designed for text-to-speech (TTS) use. It is derived by applying speech restoration to the LibriTTS corpus, which consists of 585 hours of speech data at 24 kHz
-

AlteredAvatar: Fast Style Adaptation for Dynamic 3D Avatars
By
–
AlteredAvatar: Stylizing Dynamic 3D Avatars with Fast Style Adaptation
— AK (@_akhaliq) 31 mai 2023
presents a method that can quickly adapt dynamic 3D avatars to arbitrary text descriptions of novel styles. Among existing approaches for avatar stylization, direct optimization methods can produce excellent… pic.twitter.com/k9uhlZhWz0AlteredAvatar: Stylizing Dynamic 3D Avatars with Fast Style Adaptation presents a method that can quickly adapt dynamic 3D avatars to arbitrary text descriptions of novel styles. Among existing approaches for avatar stylization, direct optimization methods can produce excellent
-

LANCE: Stress-testing Visual Models with Language-guided Counterfactual Images
By
–
LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual Images propose an automated algorithm to stress-test a trained visual model by generating language-guided counterfactual test images (LANCE). Our method leverages recent progress in large language