also the embedding-space interpolation in the GIF is very new and incredibly cool; Nomic solved the supervised version of the research problem I tweeted about here:
@jxmnop
-
Acknowledging Key Contributor to AI Model Development
By
–
finally just want to give a public shoutout to @zach_nussbaum for all the hard work he put into making this model great. it wasn't a super clear path to this level of performance and it's inspiring to see people stick with projects through that. I certainly owe him a drink
-

Nomic Embed: Open Text Embedding Model and Dataset Released
By
–
if you're doing research on text embeddings, you know that there are lots of tricks required to train a good model, and no open datasets.
— dr. jack morris (@jxmnop) 1 février 2024
we trained Nomic Embed, a great text embedding model, and actually released the data! certainly will make my research easier – check it out! https://t.co/akhYHzjKbBif you're doing research on text embeddings, you know that there are lots of tricks required to train a good model, and no open datasets. we trained Nomic Embed, a great text embedding model, and actually released the data! certainly will make my research easier – check it out!
-
Carbon Dating ML Papers by Open-Source Language Models
By
–
carbon dating ML papers by which open-source LM they use GPT-2 → the before times
GPT-J → summer 2022
LLAMA-1 → spring 2023
LLAMA-2 → summer 2023
mistral 7b → fall 2023
mixtral 8x7b → the current era -
New Embedding Model Release Announced for Tomorrow
By
–
new embedding model dropping tomorrow…… stay tuned
-
Security versus Information Theory: Different Language, Same Concept
By
–
i think it just depends on the language of your field; security people call this an "attack"; information theory people might just call this "decoding"
-

Multilingual Text Embedding Inversion Attacks on Language Models
By
–
just read a cool follow-up to vec2text: "Text Embedding Inversion Attacks on Multilingual Language Models" these folks extend embedding inversion to the *multilingual* setting, where we might not know the language of the encoded text ahead of time they add a
-
Balancing Personal Goals Against Community Impact in Tech
By
–
yeah like balancing risk and reward/impact, whether to optimize for your personal goals or those of a wider community or what the difference there is, how to navigate the market of ideas, when to know how to quit and move on to something else vs blindly keep going
-
Five Deep Topics About AI, Research and Ethics
By
–
What are five topics you can talk about for 30 minutes with zero prep 1. information content of text embeddings
2. architectural mysteries of transformer language models
3. the meta-research process (how & why to pick problems)
4. whether the NBA is rigged
5. shrimp suffering -
LLMs Achieve Success in Text Classification, Translation, and Programming
By
–
Large Language Models (LLMs) have achieved unprecedented success in tasks such as text classification (TC), machine translation (MT), computer programming (CP), and writing most sentences in the introduction of this paper (WMSITIOTP)