Multimodality was announced (well, teased with a demo) at the March debut of GPT-4. Hearing credible insider rumors that it's going to come out v soon. The Information is reporting that they're about to drop it to beat Google to the punch with Gemini, but I doubt Google is
MULTIMODAL AI
-
Multimodal GPT-4 imminent, AI image interpretation will drive AutoGPTs insane
By
–
Multimodal GPT-4 coming very soon. When AI can interpret images and web pages as good as it can interpret text, the AutoGPTs will go insane.
-

Hugging Face Releases Open-Access Multimodal AI Model
By
–
Most people missed one of the most important news of the summer – open-access multi-modal (image + text) model coming out of @huggingface
! You can optimize it, fine-tune it and customize it for your use-case: https://
huggingface.co/blog/idefics You can try the fun version of it here: -
Stable Audio Generates Music and Sound Effects from Text Prompts
By
–
With Stable Audio, you can generate music or sound effects from a text prompt and a specified duration. It uses a latent diffusion for audio model to generate audio in high-quality, 44.1 kHz stereo.
-
Imaging Tech Saves Motorcycle Rider Lives Through Safety Innovation
By
–
Not going to lie, I love this! It’ll save so many lives. Tech + imaging enabled.pic.twitter.com/COqHJAVvqZ#SafetyFirst #RoadSafety #bikes #powerbikes #motorcycles #motorbikes #Riders #tech
— Catherine Adenle (@CatherineAdenle) 19 septembre 2023Not going to lie, I love this! It’ll save so many lives. Tech + imaging enabled. #SafetyFirst #RoadSafety #bikes #powerbikes #motorcycles #motorbikes #Riders #tech
-

Exponential AI Progress Across Speech, Vision, and Language Tasks
By
–
incredible AI progress in one eye-full. this is what exponential looks like. speech, image, reading, language understanding, grade school math, codegen – all nearing or exceeded human performance
-

SeamlessM4T: Multimodal Speech Translation Model for 100 Languages
By
–
Last month we announced SeamlessM4T, a foundational multimodal model for speech translation that can perform tasks across speech-to-text, speech-to-speech and more for up to 100 languages depending on the task. More details on this work https://
bit.ly/3PqQncw -
Top CS Universities Collaborate on Computer Vision Research
By
–
.
@Stanford @UCBerkeley & @Caltech computer vision faculty& their students meet today to exchange research ideas, topics include 3D vision, language-visual models, robotic learning, computational photography, vision foundation models, etc. At the EOD, AI is truly fun Science! 1/ -
Voice Clone AI: Can You Identify the Real Voice?
By
–
Something fun with voice using generative AI! Here’re four audio clips of either my voice clone (3 clips) or me (1 clip) telling AI jokes. Can you tell which is the real me? Please reply below! Thanks @Speechlab_ai (an @AI_Fund portfolio company) for building the voice clone. pic.twitter.com/MMSPUoU2Bm
— Andrew Ng (@AndrewYNg) 18 septembre 2023Something fun with voice using generative AI! Here’re four audio clips of either my voice clone (3 clips) or me (1 clip) telling AI jokes. Can you tell which is the real me? Please reply below! Thanks @Speechlab_ai (an @AI_Fund portfolio company) for building the voice clone.
-

Bayes’ Rays: Uncertainty Quantification for Neural Radiance Fields
By
–
Bayes' Rays: Uncertainty Quantification for Neural Radiance Fields https://
bit.ly/45NTMZF
#AI #MachineLearning #DeepLearning #LLMs #DataScience