Internal performance evaluations from Google have shown significant improvements for Gemini Pro over previous Google models in understanding, summarization, reasoning, coding, and planning. This model also excels in cross-modal reasoning. (2/3)
MULTIMODAL AI
-

Gemini Pro Now Available on Poe with Vision Capabilities
By
–
Gemini Pro is now available for all users on Poe! This is the first official model on Poe with vision capabilities for image and video input. (Note: input is currently limited to web, macOS, and Windows) (1/3)
-
Runway ML Video Interpolation Technique Demonstration
By
–
Then I combined back to a video and interpolated using @runwayml. (I tried many different interpolators)
— fofr (@fofrAI) 15 décembre 2023
And here's a GIF: pic.twitter.com/0NFWnl05TpThen I combined back to a video and interpolated using @runwayml
. (I tried many different interpolators) And here's a GIF: -

User review of Pika Art as a video generation tool
By
–
No credits (for now), and fast queuing. This is probably the best AI video generator I've used so far Have you tried Pika Art already?!
-

Overview of Pika Platform Generative Video Features
By
–
The Pika platform allows you to generate a 3s video from the prompt or based on your existing media file. It also has upscaling and the possibility to extend the length of the video to 7 seconds.
-
StyleDrop: AI Model for Stylized Text-to-Image Synthesis
By
–
Introducing StyleDrop, a model that allows a significantly higher level of stylized text-to-image synthesis by using a few style reference images that describe the style for text-to-image generation, bypassing the burden of text prompt engineering. More→ https://t.co/F3Rw3QlbtP pic.twitter.com/2J4wljmFwF
— Google AI (@GoogleAI) 15 décembre 2023Introducing StyleDrop, a model that allows a significantly higher level of stylized text-to-image synthesis by using a few style reference images that describe the style for text-to-image generation, bypassing the burden of text prompt engineering. More→ https://
goo.gle/3v0PsZw -

Viewing generated images within ChatGPT voice conversations
By
–


It turns out you can see generated images right inside the voice conversation @ChatGPTapp From there you can also minimise it to the bottom corner and open back of needed. Here is a Christmas Tree generated by one of my GPTs
-

Marigold Depth Model Now Available on Replicate Platform
By
–
The excellent Marigold depth model is now on Replicate: https://
replicate.com/adirik/marigold @alaradirik for adding -
Ego-Exo4d: Large-Scale Human Activity Video Dataset
By
–
Ego-Exo4d: a large dataset of videos of people doing stuff. https://t.co/JXf209gHLb
— Yann LeCun (@ylecun) 14 décembre 2023Ego-Exo4d: a large dataset of videos of people doing stuff.
-
YouTube as AI Training Data: Audio, Video, Transcriptions
By
–
YouTube is an incredible treasure trove of data. Audio, video, transcriptions, niche communities constantly uploading…