Yeah every time we talk it's like training on each other's decoded outputs. haha.
@yitayml
-
Extreme Experiences vs Easy Examples in AI Training
By
–
I really enjoyed my casual Saturday morning coffee chat with @hwchung27
. Tons of technical wisdom and fun. With respect to life he basically said "extreme experiences are way more valuable even if they are hard" just like how training on only easy examples don't produce -
T5 Span Corruption: Data Masking and Autoregressive Decoder Loss
By
–
Span corruption in t5 is just a data operation over regular text. Masking stuff in inputs and moving them to targets. It's still fundamentally autoregressive loss on the decoder end.
-
Charformer: Google AI’s Popular Research Project Launch
By
–
@vqctran charformer yay! was one of the funnest project at @GoogleAI that got launched a disproportionate amount of times compared to our other papers lol.
-
Doubling Model Parameters Without Increased Compute Cost
By
–
~2 x the parameters for the same compute cost. Basically free model sparsity (sparse w.r.t to enc/dec blocks).
-

Encoder-Decoder vs Decoder: Clarifying AI Architecture Misconceptions
By
–
So many misconceptions about architectures (esp encoder-decoder vs decoder) partially due to nomenclature being confusing. – EncDec, PrefixLMs, Causal Dec-onlys are *all* autoregressive. Even T5/UL2's objective is autoregressive. – All 3 archs are not that different. People
-
UL2 Model Training Dataset C4 Quality Assessment
By
–
ul2 uses c4 only. c4 alone is not the best due to lack of diversity but it's pretty strong. the c4 dataset is quite good imo.