We compare CLIP (vision only) to Captioning+GPT (vision + reasoning) over 4714 images from the Emotions in Context dataset and observe a small but noticeable difference.
MULTIMODAL AI
-

CLIP Vision Model Bias: Skin Recognition and Context Understanding
By
–
Some interesting examples: In this image and a few others, CLIP appears to associate bare skin with "embarrassment". LLAVA and Captions + GPT don't, seeming to reason over the location and context.
-
Emotions Fast and Slow: Multimodal Theory of Mind Framework
By
–
New preprint on Emotions Fast and Slow! How might we perform emotional theory of mind in an image? We propose combining fast visual processing (akin to affective empathy) with slower linguistic processing (akin to cognitive empathy).
-
Multimodal AI Future: Industry Roadmap Planning Insights
By
–
Definitely changes by industry, but this is overall. Might be helpful for roadmap planning too. IMO a lot of early Applied AI will still be single modal, but multimodal AI is the future.
-
AI Processing Capabilities Ranking Text Spreadsheet Image Speech
By
–
Processing:
Text – 10
Spreadsheet – 9
Image – 6.5
Speech – 6
Code – 5 – 4
3D – 3.5
Audio – 2 -
AI Generation Capabilities Ranked by Modality Type
By
–
Generation:
Text – 10
Code – 10 – 8
Speech – 5
Image – 3
Spreadsheet – 2.5
3D – 2
Audio – 1 -
AI Brings Classic Vinyl Album Covers to Life
By
–
Las portadas de los vinilos clásicos cobran vida gracias a la inteligencia artificial y el resultado es genial.
— Juan Merodio (@juanmerodio) 2 novembre 2023
Credit: stbl_reel / r/stablediffusion pic.twitter.com/5GZFhwz3IULas portadas de los vinilos clásicos cobran vida gracias a la inteligencia artificial y el resultado es genial. Credit: stbl_reel / r/stablediffusion
-
Research@ NYC: AI Language Tech and Brain Mapping Breakthroughs
By
–
We kicked off Research@ NYC this morning with an inspiring round of lightning talks! Researchers explored important topics like using AI to extend language technologies, how we can make a map of the brain, and reconstructing music from brain activity.
-

DALL-E 3 vs DALL-E 2: Amazing Advances in Text-to-Image
By
–
(french update!) Nouvelle vidéo sur une nouvelle chaîne de vulgarisation IA 🙂 DALL·E 3 vs. DALL·E 2 : Les Avancées Incroyables en Texte à Image https://
youtu.be/pmQvn9GCkLU -

Runway ML Gen-2 Update Improves Video Generation Fidelity
By
–
We have released an update for both text to video and image to video generation with Gen-2, bringing major improvements to both the fidelity and consistency of video results.
— Runway (@runwayml) 2 novembre 2023
Try it now at https://t.co/ekldoIshdw pic.twitter.com/RyLiar7MFjWe have released an update for both text to video and image to video generation with Gen-2, bringing major improvements to both the fidelity and consistency of video results. Try it now at http://
runwayml.com
