This is a fantastic scaling effort to solve computer vision tasks. Better vision models will help us engineer better safe self-driving cars, and other types of robots
MACHINE LEARNING
-
Combining Three Transformers Achieves Strong Multimodal Image-Text Results
By
–
I love the modelling simplicity in this paper: combine 3 transformers into a big transformer and, voilà! amazing results for mapping images+text to text.
-
Progress in AI voice generation: resolved modulation, upcoming context
By
–
Update AI-Generation de voix : ⏳ (2023)
— Defend Intelligence (Anis Ayari) (@DFintelligence) 13 février 2023
Regardez c’est ouf ⬇️. Ils résolvent le problème de modulation de la voix en se mappant sur une voix en entrée.
Le nexts step ca sera de comprendre le contexte d’un texte et de générer la voix sans une autre voix en input. Stylé. https://t.co/gUf1ytr1Bu pic.twitter.com/LrSW7FF59PUpdate AI voice generation: (2023) Look, it's crazy. They solve the voice modulation problem by mapping onto an input voice. The next step will be to understand the context of a text and generate the voice without another voice as input. Cool.
-
Engineers add a line to boost publications and avoid complaints
By
–
I don't see the little line added by some engineers at the output of the recommendation model to boost these publications so that they don't come bothering them every 5 minutes.
-

Comprehensive Mathematical Foundations for Computer Science and Machine Learning
By
–
Algebra, Topology, Differential Calculus, and Optimization Theory For Computer Science and Machine Learning – Jean Gallier, UPenn A comprehensive book that covers math theories related to CS and machine learning. Very intense, 2163 pages!!! Get a copy: https://
cis.upenn.edu/~jean/math-dee
p.pdf
… -
Token Merging Cuts ViT Inference Time in Half
By
–
Token Merging can cut inference time in half & we expect it to unlock more use of large-scale ViT models in real-world applications. To enable the research community to build upon these advancements we've released the code and benchmarks here
-
Meta AI Reduces Vision Transformer Latency with Token Merging
By
–
New research from Meta AI reduces latency of existing Vision Transformer models with no additional training. ToMe combines similar tokens, reducing computation w/o losing information. Results are 2-3x speed for state-of-the-art models w/ minimal performance loss.
— AI at Meta (@AIatMeta) 13 février 2023
Read more ⬇️New research from Meta AI reduces latency of existing Vision Transformer models with no additional training. ToMe combines similar tokens, reducing computation w/o losing information. Results are 2-3x speed for state-of-the-art models w/ minimal performance loss. Read more
-

Load Dataset to Data Loader in Three Lines Code
By
–
TL;DR: You can go from a loading a dataset to a data loader in just 3 lines of code! https://
gist.github.com/Vaibhavs10/eac
93eee5b9beec6f11db7ee50eb72ea
… -
Audio Datasets Enhanced with Usage Guides on Hugging Face Hub
By
–
Your favourite Audio datasets just got whole of a lot better! 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 13 février 2023
All major audio datasets on the 🤗 Hub will now have a "How to use" page packed with example scripts, to help you get from raw data to a data loader at lightning ⚡️ speed.https://t.co/5Fzkno9DQg pic.twitter.com/RMLplHw0npYour favourite Audio datasets just got whole of a lot better! All major audio datasets on the Hub will now have a "How to use" page packed with example scripts, to help you get from raw data to a data loader at lightning speed. https://
huggingface.co/datasets/mozil
la-foundation/common_voice_11_0
…
