On the Mathematics of Diffusion Models A great resource for those who want to dive deep into mathematics of diffusion models, the kind of models that revolutionized image generation. https://
arxiv.org/abs/2301.11108
MULTIMODAL AI
-

Mathematics of Diffusion Models: Deep Dive Resource
By
–
-
Pioneer Alex Serdiuk on Building Generative AI Voice Solutions
By
–
In full here: https://
ninaschick.substack.com/p/pioneer-alex
-serdiuk-on-building
… #AI #AIVoice #GenerativeAI -
AI Voice: Performance Augmentation vs Automation and Ethics
By
–
How will AI voice augment human performance vs automate it? ‘Performance’ and the limits of AI voice cloning
Voice synthesis as a medium to expand a performer’s IP Consent and cloning of voices -

Vision-Language Models Achieve SOTA Video Classification Performance
By
–
We use pre-trained vision-language models to improve video classification, achieving SOTA performance on Kinetics-400 and surpassing previous methods by 20-50% on five popular video datasets. #AAAI2023 Github: https://
github.com/whwu95/Text4Vis -

AI-Generated Image of Henry Ford’s Autonomous Car Design
By
–
"Autonomous Car designed by Henry Ford" #StableDiffusion2 #AIart
-

ControlNet and Stable Diffusion pose control experimentation
By
–
Juguemos con ControlNet y Stable Diffusion. La imagen de input (yo) servirá para darle una pose de control al output, y no tiene por objetivo que se parezca físicamente a mi. PROMPT: An old man
-
Robots Learn Natural Language Communication and Real-World Skills
By
–
Part 6 in the series is out, discussing our work in robotics, especially allowing robots & humans to communicate naturally via language, apply common sense knowledge in real-world situations & increasing the number of low-level skills robots can perform.
-
Unimodal Text Models Unfairly Compared in Visual QA Tasks
By
–
Lo vi y lo estuve estudiando meter en el vídeo PEEEEEERO tiene un poco de trampa. Comparar modelos unimodales de texto (ciegos como GPT 3.5) en tareas de QyA visuales es un poco raro. Y más teniendo la opción de validarlo contra BLIP.
-

BLIP-2: ChatGPT with Image Vision Capabilities
By
–
¡¡NUEVO VÍDEO!! ¿Y si ChatGPT pudiera VER una imagen? Es una idea muy sugerente, ya que poder dialogar sobre el contenido de una IMAGEN amplía mucho el rango de tareas que poder hacer. Pues con BLIP-2 algo así es posible. Te lo cuento! https://
youtu.be/VZQJuY-cL8w
ㅤ -

ControlNet: Fine-Grained Control for AI Image Generation
By
–
De forma resumida, ControlNet es una de las piezas faltantes en las IAs de generación de imágenes, donde ahora puedes tener un mayor control, usando como inputs mapas de profundidad, bordes, poses, etc. Ahora sí, es control puro al proceso de generación.