Multilingual LongT5 that uses UL2 objective.
@yitayml
-
Base Models Achieve In-Context Learning Through Instruction Tuning
By
–
i think even <base models have in-context learning ability if instruction tuned.
-

Comparing Language Models: Sparse, Dense, and Parameter Scaling Issues
By
–
so many problems i don't know where to begin.
– yea put sparse and dense models in the same plot with the # params. good job – i'm sure you know the size of palm-2 and gpt-4. – fwiw, t5 is still one of the best LM models out there. it started way earlier than 2021. – -
Real URLs are hard: identifier stability challenges
By
–
yeah. real urls are hard. really hard. and especially if identifiers change over time. investigated this so much that we know it doesn't actually work in the wild.
-

Symbol Tuning Improves In-Context Learning Research
By
–
Once again, another piece of great science from @JerryWeiAI that shows that symbol tuning improves in-context learning. Check out his thread below
-
Fine-tuning Flan collection with T5X and SeqIO
By
–
put the flan collection into seqio and finetune it in t5x? even i don't have access to this code anymore. im not sure if the seqio task registries have been open sourced.
-
FLAN-T5 and FLAN-UL2 Outperform Vanilla Models
By
–
they are orthogonal concepts. flan-{t5/ul2} is almost always better than the vanilla model.
-
UL2 Fine-tuning Outperforms T5 in Most Cases
By
–
Fwiw, fine-tuning ul2 is almost always better than t5.
-
540 Billion Parameters: Scaling Language Models
By
–
Hahahah, 540 is already a reserved number in my head that if you prompt me with "540" I will continue by saying "billion parameters" almost instinctively.
-
Research Integrity Over Hype in AI Model Distillation
By
–
I enjoyed the rigour and research flair of this work. In particular, the "self-respect" to do the right thing to advance science and not hype grab by obfuscating model distillation as progress. This is what the community should expect from a top research institute.
