New piece on emergence in language models by @JacobSteinhardt: https://bounded-regret.ghost.io/emergent-deception-optimization/#fnref7 I found the takeaways quite lucid:
– Capabilities that would lower training loss will emerge in the future
– As models scale up, simple heuristics tend to get replaced by complex ones
MACHINE LEARNING
-
Emergence in Language Models: Capabilities and Heuristics
By
–
-
Retrieval-Augmented Language Models Discussion with MLStreetTalk
By
–
Tune in to an electrifying discussion on retrieval-augmented language models with @MLStreetTalk
, @PSH_Lewis
, and the adorable Cohere mascot, Mable! : -
Daily Python Machine Learning and Language Models Tutorials
By
–
Every day, I share tutorials and simplify complex topics around Python, Machine Learning & Language Models. Follow me → @Sumanth_077 to ensure you don't miss that. Like/RT the first tweet to support my work and help this reach more people.
-
Abacus AI Praised for Building End-to-End ML Systems
By
–
Absolutely Abacus AI is really great to build End To End ML systems
-
Flan-T5 vs GPT-3.5: Fine-tuning and Zero-shot Comparison
By
–
Great post! Just a clarification, was Flan-T5 further finetuned on any data or was it both few/zero shot for gpt3.5 and Flan-T5?
-
Few Organizations Can Train 100B+ Parameter Models Estimate
By
–
Definitely way less than 200. A wide spectrum on what it means to "train 100B+ parameter models". But I would estimate this number to be <50 optimistically.
-
Instruction Tuning as a Form of Data Improvement
By
–
Yeah, instruction tuning is kinda like "better data"
-
Finetuning vs Emergence: Different Paths to Model Abilities
By
–
Finetuning is a great way to bring an ability to a smaller model. Emergence often occurs for an ability that is *not* explicitly trained for, just by increasing scale (compute/ model size).