those too, but i think more importantly i'm asking about the relationship between model parameters, training steps (whatever that means in the embeddings case), and downstream performance
@jxmnop
-
Scaling Laws for Language Models Explained
By
–
no, scaling laws like the scaling laws for language models:
-
Scaling Laws for Embedding Models: Research Initiatives
By
–
who's working on scaling laws for embedding models? could be any of:
• text embeddings (DPR, GTR, GTE…)
• image embeddings (SimCLR, DINO…)
• recommendation systems (?)
• multimodal embeddings (CLIP, ImageBind…)
• any other type of embeddings… -
Nomic AI’s latest developments in open source language models
By
–
literally @nomic_ai (cc @andriy_mulyar
) -
HuggingFace Model Quantization on 24GB GPU
By
–
I used HuggingFace and quantized the model to 4 bits, I believe it ran on a 24GB gpu (maybe need 48)
-
Circumventing Biden AI Model Training Restrictions Through Parameter Scaling
By
–
the biden executive order put restrictions on models that:
• contain at least 10^9 parameters
• use more than 10^26 floating-point operations my future company will get around these restrictions by simply training a 9,999,999 billion param model comprised of 8192-bit floats -
Circumventing Biden’s AI Model Parameter and Compute Restrictions
By
–
the biden executive order put restrictions on models that:
• contain at least 10^9 parameters
• use more than 10^26 floating-point operations my future company will get around these restrictions by simply training a 9,999,999 billion param model comprised of 8192-bit floats -
4096 Bit Floats: Advanced Computing Technology Innovation
By
–
shhhh, this was gonna be how i won the bet, 4096 bit floats
-
Measuring representational capacity in AI models
By
–
but why not? how do you measure representational capacity?
-
Timeline for GPT3 and GPT4 Level AI Models Development
By
–
5 years for a gpt3-level 1B model, 8 years for gpt4-level