this assumes we can’t move the Pareto curve of scaling laws
@jxmnop
-
Can a 1B Parameter Model Ever Surpass GPT-4 Capabilities
By
–
will there ever be a 1B param model that’s more capable in every way than GPT-4 is today?
-
Training Costs of AI Models: OpenAI GTR and SBERT
By
–
yeah mostly bc they’re expensive to train, rn we have openAI GTR and sbert
-
Revolutionary vec2text Library Achieves Perfect Vector-to-Text Conversion
By
–
yes!! I’m actually the only person in the world rn with a system that can do this perfectly. here’s the library: https://
github.com/jxmorris12/vec
2text
… -
Word2Vec Paper Relevance Assessment for ML Research
By
–
don’t think this paper on word2vec is relevant to my work at all!
-
Inferring LM Prompts from Output Probabilities Without API Access
By
–
haven't posted anything about (2) yet but here's the TLDR: • given LM output probabilities, we can infer what the input prompt was
• we built a model that can do this
• most APIs don't give you probabilities, but we came up with a clever algorithm to get them using logit bias -
vec2text: Reconstructing Text from Embeddings Paper
By
–
more on this second paper soon but the paper draft is here http://
openreview.net/forum?id=t9dWH
pGkPj
… (hopefully w/ a new version on arxiv later this week) and all the code for both papers is available in vec2text: https://
github.com/jxmorris12/vec
2text
… -

Language Model Inversion Research Talk by Sasha
By
–
Sasha gave this amazing talk on our language model inversion research! 1. text embedding inversion (
http://
arxiv.org/abs/2310.06816)
2. language model output inversion (coming soon…) -
Latent space sequence-to-sequence translation advantages explained
By
–
what do you think are the advantages of doing this kind of sequence-to-sequence translation in latent space?
-
Text Diffusion Research: Current State and Paper Critique
By
–
that's funny! but yeah, there's a lot of work on text diffusion already! I don't think this paper is a representative or useful example!