Experience the next level of chatbot innovation with @Microsoft
's Kosmos 1, elevating ChatGPT with improved voice and visual commands. Discover how AI is revolutionising user experiences with this amazing technology here: https://
bit.ly/3FD4hUX @OpenAI @jeevprabnivash
MULTIMODAL AI
-
Microsoft Kosmos 1 Elevates ChatGPT with Voice and Visual Commands
By
–
-

Microsoft hints at multimodal GPT-4 and Visual ChatGPT
By
–
Yesterday, Microsoft Germany CTO dropped a comment about upcoming GPT-4 being multimodal (handling images and text) But the day before, Microsoft published "Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models" Abstract of the https://
arxiv.org/pdf/2303.04671
.pdf
… -
Stacking APIs: Face, speech cloning and GPT replicate conversations
By
–
Probably 90%+. Was waiting for someone to stack these APIs. The (Face+speech cloning+GPT) proof of concept is the precursor to replicating conversations with anyone. Quality is not jaw-dropping yet, but it will be soon.
-
D-ID combines AI presenters with ChatGPT for conversational interface
By
–
D-ID gave ChatGPT a face:
— AI Breakfast (@AiBreakfast) 10 mars 2023
The company combined their AI presenters with ChatGPT for a conversational interface
Free to try with account: https://t.co/rKPtFYEbzQ
🍌 pic.twitter.com/40Byz3pH6uD-ID gave ChatGPT a face: The company combined their AI presenters with ChatGPT for a conversational interface Free to try with account: https://
chat.d-id.com -
PaLM-E: 562B Parameter Embodied Language Model for Robotics
By
–
Today we share PaLM-E, a generalist, embodied language model for robotics. The largest instantiation, 562 billion parameters, is also a state-of-the-art visual-language model, has PaLM’s language skills, and can be successfully applied across robot types →https://t.co/iyaPdN4SfV pic.twitter.com/foXnzgZnWd
— Google AI (@GoogleAI) 10 mars 2023Today we share PaLM-E, a generalist, embodied language model for robotics. The largest instantiation, 562 billion parameters, is also a state-of-the-art visual-language model, has PaLM’s language skills, and can be successfully applied across robot types →
https://
goo.gle/3JsszmK -

Data2vec 2.0: 16x Faster Self-Supervised Learning for Vision, Speech, Text
By
–
Created by Meta AI researchers, Data2vec 2.0 can train self-supervised models for vision, speech and text up to 16x faster than the most popular existing algorithm for images — achieving the same accuracy. Read more & access the code
-
Insert New Face into Template Using Stable Diffusion
By
–
5/ Insert a new face into a template using Stable Diffusion.
-

Computer Vision Pose Estimation Optimizes Smart City Mobility
By
–
How #ComputerVision-Powered Pose Estimation Can Optimize #SmartCity #Mobility
by @joshinav @BBNTimes_en Go to: https://
buff.ly/3yjG02R #AI #IoT #BigData #MachineLearning #ArtificialIntelligence #ML #MI cc: @bigdata @patrickgunz_ch @amuellerml -
Typeface Raises 65 Million Dollars for AI Platform
By
–
Typeface raises 65 million dollars to develop its platform trained on ChatGPT and Stable Diffusion models https://actuia.com/actualite/typeface-leve-65-millions-de-dollars-pour-developper-sa-plateforme-formee-sur-les-modeles-chatgpt-et-stable-diffusion/
… #AI #artificialintelligence #fundinground -
Papers with Code: 7000+ Public Datasets Across Multiple Modalities
By
–
3. Papers with Code Papers with Code consist of more than 7000 Public Datasets on different modalities. Find them here: