Research from Meta AI reduces latency of existing Vision Transformer models with no additional training. Token Merging can cut inference time in half and we expect it to unlock more use of large-scale ViT models in real-world applications. Read more https://
bit.ly/3ZJv61D
MULTIMODAL AI
-
Meta AI Reduces Vision Transformer Latency with Token Merging
By
–
-
Foundational Model Generates Video and Text Outputs
By
–
This is not just a plain old text-generating model — but it generates VIDEO (from text prompts) too. This would make it the first foundational general-purpose model to generate outputs across different mediums (text and video, in this case.)
-
Midjourney Update: Exploring Technology and Artistic Creativity
By
–
In this video, we check out Midjourney’s update, which shows the artist's creative process and showcases some of their latest creations. This project offers a fresh perspective on the relationship between technology and creativity.https://t.co/TkuCniu6nv
— Satya Mallick (@LearnOpenCV) 13 mars 2023
We invite viewers to… pic.twitter.com/2vMCPAgd69In this video, we check out Midjourney’s update, which shows the artist's creative process and showcases some of their latest creations. This project offers a fresh perspective on the relationship between technology and creativity. https://
youtube.com/watch?v=L0NBQz
y39EQ
… We invite viewers to -
Multi-modal AI and the fragility of Stable Diffusion prompting
By
–
yeah entering multi-modal territory right there the entire "art" of stable diffusion prompting could def be wiped out next week
-

Conversational AI Building Experiential Foundation in Metaverse
By
–
Into the metaverse: How conversational #AI will build its experiential foundation
by @rajkoneru @VentureBeat Go to: https://
buff.ly/3xodMTt #Digital #BigData #ArtificialIntelligence #MI #Blockchain cc: @wil_bielert @patrickgunz_ch @terenceleungsf -

GigaGAN: Scaling GANs for Fast High-Resolution Image Synthesis
By
–
10/ GigaGAN – enables scaling up GANs on large datasets for text-to-image synthesis; it’s found to be orders of magnitude faster at inference time, synthesizes high-resolution images, & supports various latent space editing applications.
-

Visual ChatGPT Enables Multimodal Interaction Beyond Text
By
–
3/ Visual ChatGPT – it connects ChatGPT and different visual foundation models to enable users to interact with ChatGPT beyond language format.
-

Tech Journalist and AI Marketing Director Profile
By
–
#NewProfilePic تحياتي ومحبتي إعلامي تكنولوجيامدير تسويق وعلاقات لمجموعة عالميةصانع محتوىدبلوم صحافة ذكاء اصطناعي وميتافرس
Tech JournalistGroup Marketing & PR MngrAI & Meta -

Visual ChatGPT: Connecting Multiple Foundation Models via Prompting
By
–
I'm super impressed by the "visual" ChatGPT by Microsoft Research Asia, connecting ChatGPT with more than dozen of open-source visual/captioning foundation models through smart prompting
— Thomas Wolf (@Thom_Wolf) 11 mars 2023
Code is super simple, it's a single python file built on top of https://t.co/MQ5PcVyjGi… pic.twitter.com/spcIohmOEfI'm super impressed by the "visual" ChatGPT by Microsoft Research Asia, connecting ChatGPT with more than dozen of open-source visual/captioning foundation models through smart prompting Code is super simple, it's a single python file built on top of https://
github.com/microsoft/visu
al-chatgpt
… -
Disney Imagineers Develop Emotional Connection Robots at SXSW
By
–
As a little girl I wanted to be a #Disney Imagineer. Disney at #SXSW now: "this is our latest effort in making robots that can make an emotional connection with our guests" https://
youtube.com/clip/Ugkxx6y3V
ptridTbOj-VRRMEiZHVQXbJ44E1
…