Also thanks to people who made amazing colabs and ported it to @huggingface at such lightning speed!
@yitayml
-
Flan-UL2 Model Port Already Available
By
–
Maybe u should port Flan-UL2 too! Oh wait… Already there.
-
Flan2 Missing from Rankings Despite Late Release
By
–
And Flan2 should have been there despite coming out only the last quarter. @ZetaVector is still figuring out why it's not there. But citations are just for fun amirite?
-
Chain-of-Thought and Self-Consistency Beat Direct Prompting
By
–
I think this is more about the task than the models. This is also the same for flan palm 62b and even 540b (except for BBH IIRC). That said, cot + self consistency is usually better than direct prompting! I go over this a little in the blogpost.
-
Scaling Strategy Over Fine-Tuning for Emergence
By
–
Probably should also leave them a note to tell them to scale up for emergence and not wasting a couple of years on fine-tuning 100m encoder models.
-
FLAN2 Paper on Language Model Fine-tuning Released
By
–
Oh I mean the flan2 paper. https://
arxiv.org/abs/2210.11416 -
Flan and UL2 Missing from Citations List
By
–
Hmm. Where is Flan and UL2 on the list? They have more than 40 citations for sure. And probably more. IIRC flan2 has much more than that.
-
C4 Pretraining Limits in Base Model Development
By
–
We might be hitting limits of the c4 pretraining that the base model uses as well.
-
Flan-PaLM 62B Model Shows Modest Performance Improvements
By
–
Good point. It's a modest boost. When you take into account Flan-PaLM 62b results, this falls within expectation. It's modest. We didn't expect game changing results from this model either. We just give folks the option to squeeze a little more quality if they could afford to.
-
UL2 Model Now Available on Hugging Face
By
–
The ul2 model is already on Huggingface. I don't know how HF works but it should be minimal work to make it available there. It's an identical model config to ul2 20b. Cc:
@huggingface