This is kind of weird, if you think about it. The unknown words together will be more frequent than the 9998th most frequent term. So the fidelity of representation isn't being distributed well. How can we give the unknowns more vectors?
MACHINE LEARNING
-
Word representations learning for smaller datasets and per-token decisions
By
–
For language modelling and translation, word pieces are the standard way to address this. But for smaller datasets, especially where the decisions need to be made on a per-token basis, it's very helpful to learn representations for words, rather than characters or word pieces.
-
spaCy Trained Pipelines: Domain Flexibility and Fine-tuning Capabilities
By
–
But for spaCy, we want to give people trained pipelines and have them be as applicable as possible across different domains. We also want people to be able to fine-tune the models on their data if it's not working ideally.
-

The Hashing Trick in spaCy Models: An Underexplored Technique
By
–
The hashing trick is an old technique from sparse linear models, well known in toolkits like Vowpal Wabbit. It's one of the unusual things I did in @spacy_io 's models that I've always felt was quite neat, and needed more experiments and write up.
-
Unbounded Vocabularies and Fixed-Size Embedding Tables Explained
By
–
The basic motivation is that vocabularies are unbounded, but embedding tables can only be a fixed size. If you know your training data matches up to your test data well, this isn't such a big deal for most applications. If it's rare at training time, it'll be rare at test time.
-
ChatGPT: Millions of Linguists Working on Language Model
By
–
Imagine millions of linguists working for decades on an extremely complex model of language based on all the text ever written. That’s essentially what ChatGPT is.
-

AI-Based Fitness Trainer Application for Squat Analysis
By
–
Today we will build an AI-based Fitness Trainer application that analyzes squats and provides appropriate feedbackhttps://t.co/6987UxxwwK#fitnesstrainer #squatanalyzer #squat #artificialintelligence #computervision #deeplearning #ai #machinelearning #poseestimation #mediapipe pic.twitter.com/64cWBTP5Rx
— Satya Mallick (@LearnOpenCV) 20 décembre 2022Today we will build an AI-based Fitness Trainer application that analyzes squats and provides appropriate feedback https://
learnopencv.com/ai-fitness-tra
iner-using-mediapipe/
… #fitnesstrainer #squatanalyzer #squat #artificialintelligence #computervision #deeplearning #ai #machinelearning #poseestimation #mediapipe -
Annotation Quality Comparison and ML Impact on Professional Fields
By
–
Does the paper also compare the quality of the annotations? Re: Tsunami. There was a recent thread about how the market for translation was destroyed by ML. It will happen one field at a time… so the impact is incremental.
-

Trustworthy ML Workshop at ICLR 2023: Statistical Computational Limitations
By
–
A super cool and timely workshop to be co-hosted at #ICLR2023! When can statistical and computational limitations arise in the context of trustworth ML? Deadline is February 8, check out their unique two track system. https://
sites.google.com/view/trustml-u
nlimited/call-for-papers
… Look forward to being there! -
Software Component Requires Statistics and Rules, Not Modern AI
By
–
Yes, a software component is necessary but not AI in the modern sense. Basic statistics, basic rule-based system.