Imagine we gave the unknowns the last 100 rows instead. When we see an unknown word, we'll hash its word form, and mod the result to pick a row for it. Each unknown word will share its vector with lots of others, but its vector-mates are probably rare.
RESEARCH
-
Hashing Trick for Fixed-Size Embedding Tables in Fine-Tuning
By
–
Resizing the embedding table for fine-tuning is pretty awkward. We'd rather have some way of making sure that novel words can get a unique representation, even from a fixed-size embedding table. The hashing trick achieves this.
-
Improving Vector Representation for Unknown Words in NLP
By
–
This is kind of weird, if you think about it. The unknown words together will be more frequent than the 9998th most frequent term. So the fidelity of representation isn't being distributed well. How can we give the unknowns more vectors?
-
Word representations learning for smaller datasets and per-token decisions
By
–
For language modelling and translation, word pieces are the standard way to address this. But for smaller datasets, especially where the decisions need to be made on a per-token basis, it's very helpful to learn representations for words, rather than characters or word pieces.
-

The Hashing Trick in spaCy Models: An Underexplored Technique
By
–
The hashing trick is an old technique from sparse linear models, well known in toolkits like Vowpal Wabbit. It's one of the unusual things I did in @spacy_io 's models that I've always felt was quite neat, and needed more experiments and write up.
-
Unbounded Vocabularies and Fixed-Size Embedding Tables Explained
By
–
The basic motivation is that vocabularies are unbounded, but embedding tables can only be a fixed size. If you know your training data matches up to your test data well, this isn't such a big deal for most applications. If it's rare at training time, it'll be rare at test time.
-
Annotation Quality Comparison and ML Impact on Professional Fields
By
–
Does the paper also compare the quality of the annotations? Re: Tsunami. There was a recent thread about how the market for translation was destroyed by ML. It will happen one field at a time… so the impact is incremental.
-

Petals Creates Free Distributed Network for Running Text-Generating AI
By
–
Cool article from @Kyle_L_Wiggers @TechCrunch about Petals by @BigscienceW https://
techcrunch.com/2022/12/20/pet
als-is-creating-a-free-distributed-network-for-running-text-generating-ai
… -

Trustworthy ML Workshop at ICLR 2023: Statistical Computational Limitations
By
–
A super cool and timely workshop to be co-hosted at #ICLR2023! When can statistical and computational limitations arise in the context of trustworth ML? Deadline is February 8, check out their unique two track system. https://
sites.google.com/view/trustml-u
nlimited/call-for-papers
… Look forward to being there! -
AI Ethics and Software: A 40-Year Delay Pattern
By
–
You heard me say "AI Ethics is just software Ethics 40 years too late." "AI is just software 40 years too late!"