My iPhone puts a data grid on my face ten times every second.
MULTIMODAL AI
-
Hando Plans AI-Powered Language Learning Integration Expansion
By
–
In the future, Hando aims to integrate @CampLingo even more into its users' daily lives. He plans on releasing: Infinite stories generated by AI
Custom models for low-resource languages
Custom integration with Siri
Implementing sign languages -

SQuId: AI Model for Speech Quality Assessment Across Languages
By
–
Learn how SQuId (Speech Quality Identification), a 600M parameter regression model that describes to what extent a piece of speech sounds natural, can be used to complement human ratings for the text-to-speech evaluation of many languages → https://
goo.gle/3J1kghg -

DataComp-1B CLIP Model Outperforms OpenAI with Lower Compute Cost
By
–
Delve into DataComp with @lschmidt3 next at https://
future.snorkel.ai about innovating #ML by experimenting with new training sets. A new CLIP model, DataComp-1B, outperforms OpenAI’s by 3.7pp on ImageNet, and a 9x improvement in compute cost vs. LAION-5B. -
Complete Audio AI Course: Basics to Speech Processing
By
–
We've made this course to help you get up and running with all things Audio. From basics of audio to classification to speech-to-text to text-to-speech. This self-paced course will cover it all!
-
Japanese speaker translates English emotions into acoustic equivalents
By
–
This Japanese speaker translates English emotions into their Japanese acoustic or multimodal equivalents and it's fascinating
-
Neuralangelo: NVIDIA’s New AI Model for 3D Reconstruction
By
–
Deep learning: Neuralangelo, the new AI model from @nvidia Research for 3D reconstruction https://actuia.com/actualite/deep-learning-neuralangelo-le-nouveau-modele-dia-de-nvidia-research-pour-la-reconstruction-3d/
… #AI #artificialintelligence #deeplearning -

Recognize Anything: Strong Foundation Model for Image Tagging
By
–
Recognize Anything: A Strong Image Tagging Model paper page: https://
huggingface.co/papers/2306.03
514
…
demo: https://
huggingface.co/spaces/xinyu12
05/Tag2Text
… present the Recognize Anything Model (RAM): a strong foundation model for image tagging. RAM can recognize any common category with high accuracy. RAM introduces a -

Grounding Instructional Steps in Narrated How-To Videos
By
–
Learning to Ground Instructional Articles in Videos through Narrations paper page: https://
huggingface.co/papers/2306.03
802
… present an approach for localizing steps of procedural activities in narrated how-to videos. To deal with the scarcity of labeled data at scale, we source the step -

Emergent Correspondence Discovery in Image Diffusion Models
By
–
Emergent Correspondence from Image Diffusion paper page: https://
huggingface.co/papers/2306.03
881
… Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We