Augmenting CLIP with Improved Visio-Linguistic Reasoning paper page: https://
huggingface.co/papers/2307.09
233
… Image-text contrastive models such as CLIP are useful for a variety of downstream applications including zero-shot classification, image-text retrieval and transfer learning. However,
MULTIMODAL AI
-

Augmenting CLIP with Improved Visio-Linguistic Reasoning
By
–
-

DS-Fusion: Generating Artistic Typography with Discriminated Stylized Diffusion
By
–
DS-Fusion: Artistic Typography via Discriminated and Stylized Diffusion paper page: https://
huggingface.co/papers/2303.09
604
… introduce a novel method to automatically generate an artistic typography by stylizing one or more letter fonts to visually convey the semantics of an input word, while -
OpenAI’s Public Opinion Approach to AI Recognition Features Questioned
By
–
OpenAI says it will follow public opinion on whether GPT4 should include face, emotion and gender recognition. Which public? In which nations? Whose opinions will count?
-

Explore Your Alter Egos with Gen-1 AI Tool
By
–
Explore your alter egos with Gen-1. pic.twitter.com/6nTHeswEfa
— Runway (@runwayml) 18 juillet 2023Explore your alter egos with Gen-1.
-

INVE: Real-time Interactive Neural Video Editing Solution
By
–
INVE: Interactive Neural Video Editing
— AK (@_akhaliq) 18 juillet 2023
paper page: https://t.co/2TqK9e8eqF
present Interactive Neural Video Editing (INVE), a real-time video editing solution, which can assist the video editing process by consistently propagating sparse frame edits to the entire video clip.… pic.twitter.com/J9mesLtMA7INVE: Interactive Neural Editing paper page: https://
huggingface.co/papers/2307.07
663
… present Interactive Neural Editing (INVE), a real-time video editing solution, which can assist the video editing process by consistently propagating sparse frame edits to the entire video clip. -

Language Conditioned Traffic Generation for Autonomous Driving Simulation
By
–
Language Conditioned Traffic Generation paper page: https://
huggingface.co/papers/2307.07
947
… Simulation forms the backbone of modern self-driving development. Simulators help develop, test, and improve driving systems without putting humans, vehicles, or their environment at risk. However, -

BuboGPT: Visual Grounding in Multi-Modal Large Language Models
By
–
BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
— AK (@_akhaliq) 18 juillet 2023
paper page: https://t.co/e6caibcFrJ
LLMs have demonstrated remarkable abilities at interacting with humans through language, especially with the usage of instruction-following data. Recent advancements in LLMs, such as… pic.twitter.com/QW4crgX9fyBuboGPT: Enabling Visual Grounding in Multi-Modal LLMs paper page: https://
huggingface.co/papers/2307.08
581
… LLMs have demonstrated remarkable abilities at interacting with humans through language, especially with the usage of instruction-following data. Recent advancements in LLMs, such as -
Exploring InstructPix2Pix: AI-Powered Image Editing Technology
By
–
🎥 Dive into the fascinating world of InstructPix2Pix! https://t.co/hUEdCQnci9
— Satya Mallick (@LearnOpenCV) 17 juillet 2023
Watch as it unveils mind-blowing creations and unlocks endless potential. Join us on this journey to discover and unleash the magic of #InstructPix2Pix right at your fingertips! #AI #computervision pic.twitter.com/CP5y9uUbWvDive into the fascinating world of InstructPix2Pix! https://
youtube.com/watch?v=RsCmJo
Hgzw0
… Watch as it unveils mind-blowing creations and unlocks endless potential. Join us on this journey to discover and unleash the magic of #InstructPix2Pix right at your fingertips! #AI #computervision -
Perceiver IO: Unified Neural Network for Multiple Modalities
By
–
"Life would be drastically simpler if a single neural network architecture could handle a wide variety of both input modalities and output tasks." – Perceiver IO
-

NIFTY: Neural Fields for Human-Object Interaction Motion Synthesis
By
–
NIFTY: Neural Object Interaction Fields for Guided Human Motion Synthesis paper page: https://
huggingface.co/papers/2307.07
511
… address the problem of generating realistic 3D motions of humans interacting with objects in a scene. Our key idea is to create a neural interaction field attached to a