ICYMI: ChatGPT now can show images in voice UI and has GPT starters on iOS https:// buff.ly/3Rjy7Ti
MULTIMODAL AI
-
Brown Researchers Develop Natural Language Robot Control System
By
–
researchers from Brown work on breaking human-robot barrier: new system translates natural language into robot actions
-
AI Voice Transcription Advancements: Distill Whisper Enhances Interactions
By
–
Explore how AI voice transcription advancements, like Distill Whisper, can enhance AI interactions. See it in action in this video:
-
AI Voice Interaction Evolution: Whisper and Distill Whisper Advances
By
–
AI voice interaction is evolving to match text communication. @OpenAI
's Whisper converts voice to text effectively, and Distill Whisper by @huggingface now offers similar accuracy with increased speed and less complexity. -
Distill Whisper: 99% Accuracy with 2% Data for Real-Time Voice AI
By
–
Distill Whisper uses knowledge distillation to train a smaller model to achieve 99% accuracy of a larger one with 2% data, aiming for real-time AI voice applications without delays.
-
Gaussian-SLAM: Neural RGBD Method for Photorealistic Scene Reconstruction
By
–
8/ Gaussian-SLAM – a neural RGBD SLAM method capable of photorealistically reconstructing real-world scenes without compromising speed and efficiency.https://t.co/HqggU9MvOV
— DAIR.AI (@dair_ai) 17 décembre 20238/ Gaussian-SLAM – a neural RGBD SLAM method capable of photorealistically reconstructing real-world scenes without compromising speed and efficiency.
-
Audiobox: Unified Flow-Matching Model for Audio Generation
By
–
3/ Audiobox – a unified model based on flow-matching capable of generating various audio modalities; designs description-based and example-based prompting to enhance controllability and unify speech and sound generation paradigms.https://t.co/OcXaDuRU6j
— DAIR.AI (@dair_ai) 17 décembre 20233/ Audiobox – a unified model based on flow-matching capable of generating various audio modalities; designs description-based and example-based prompting to enhance controllability and unify speech and sound generation paradigms.
-
Pika Labs launches groundbreaking AI video generation features
By
–
ICYMI: Pika Labs rolls out groundbreaking AI video generation features https://buff.ly/47WUkxu
-
Handwriting to Text GPT converts note images to digital text
By
–
You can't keep track of all your handwritten notes.
— God of Prompt (@godofprompt) 16 décembre 2023
That's why I made Handwriting to Text GPT.
Simply upload any image of your notes/sketches/documents and it will transcribe them for you in seconds!
Try it out here: https://t.co/1CIaWAQcgi#handwriting #Text #customgpts… pic.twitter.com/g2lELOeqFYYou can't keep track of all your handwritten notes. That's why I made Handwriting to Text GPT. Simply upload any image of your notes/sketches/documents and it will transcribe them for you in seconds! Try it out here: https://
godofprompt.ai/gpts/handwriti
ng-to-text-gpt
… #handwriting #Text #customgpts -
Gemini Pro Expands Multimodal Support Across Poe Platform
By
–
The Gemini Pro bot currently supports text, image and video input with text output. We will introduce additional multimodal support in the near future. You can try it now at https://
poe.com/Gemini-Pro and across Poe apps on all platforms. Enjoy! (3/3)