Everyone is working on multimodal systems.
The question is how to do it.
And the problem is that the kind of generative architecture that works for text does not work for images and video.
MULTIMODAL AI
-
Multimodal AI: Architectural Challenges Beyond Text Generation
By
–
-
ChatGPT 4 Vision Capabilities Demonstrate Advanced AI Progress
By
–
ChatGPT 4 Vision capabilities seem like magic!
-

Latent Consistency Models for Fast Video-to-Video Conversion
By
–
Using a latent consistency model for video2video is fast, but it needs control mechanisms.
— fofr (@fofrAI) 28 octobre 2023
The speed means you can do high frame rate video conversions. But the lack of control makes it a mess.
180 frames in 55 seconds:https://t.co/Y13KKTdAtp pic.twitter.com/CMaWM62C9AUsing a latent consistency model for video2video is fast, but it needs control mechanisms. The speed means you can do high frame rate video conversions. But the lack of control makes it a mess. 180 frames in 55 seconds: https://
replicate.com/p/qavcahlbjto6
6zcoq3eyubaclu
… -
Dream Gaussian and AnimateDiff: 3D Logo Video Generation
By
–
An early experiment with:
— fofr (@fofrAI) 27 octobre 2023
– Replicate logo image to 3d video with Dream Gaussian
– Video used as QR controlnet for animatediff
– Frame interpolated using ST-MFNet pic.twitter.com/2b5fdPVLmsAn early experiment with: – Replicate logo image to 3d video with Dream Gaussian
– used as QR controlnet for animatediff
– Frame interpolated using ST-MFNet -
ChatGPT Voice as Long-Form Conversational Brainstorming Partner
By
–
ChatGPT Voice as a long-form conversational brainstorming partner:
-
Can artificial intelligence see like human beings?
By
–
#Opinion
Can artificial intelligence see like human beings? by @SophieFayad92, Doctor in Neurosciences, Research Center @talan_fr https://actuia.com/contribution/sophie-fayad/lintelligence-artificielle-peut-elle-voir-comme-les-etres-humains/
… -
ViT and ConvNets Achieve Equal Performance at Same Compute
By
–
Compute is all you need.
For a given amount of compute, ViT and ConvNets perform the same. Quote from this DeepMind article: "Although the success of ViTs in computer vision is extremely impressive, in our view there is no strong evidence to suggest that pre-trained ViTs -

Midjourney’s Peak Generation Capabilities Impress Users
By
–
What’s more impressive in my opinion here is that we can to specific peaks. 1. Antelao 2. Kazbek
3. Monviso Midjourney is so powerful it’s almost impossible to stop playing with it. -
Google Researchers Showcase AI Innovations at São Paulo Event
By
–
That’s a wrap on an inspiring day of lightning talks and interactive demos at Research@ São Paulo! Google Researchers explored topics like Enhancing Foundation Models, Med-PaLM, Dermatology On Lens, Floods, Weather Forecasting with ML, and more.
