New MC4 refreshed corpus and upgraded mT5 checkpoints from @GoogleAI
.
@yitayml
-
Google releases refreshed MC4 corpus and upgraded mT5 checkpoints
By
–
-
UL2 Pretraining Code Found in Megatron-LM Repository
By
–
day 20+ of exploring the wilderness: i found ul2 pretraining code in megatron LM. https://
github.com/NVIDIA/Megatro
n-LM/pull/268
… a nice surprise. -
Crunchbase LLM Hallucination Issues and Data Accuracy
By
–
Maybe crunchbase's LLM hallucinated that info…
-

Google AI Releases Improved MC4 Corpus and uMT5 Models
By
–
Sharing a piece of work I contributed to while at @GoogleAI
: * a new improved Mc4 corpus (29T char tokens and 107 languages) that gets language sampling right with UniMax sampling. * open source pretrained uMT5 models trained on 1T tokens. * Unimax sampling solves some -

Reverse Instructions Method Naming Discussion Paper
By
–
Its nice to see someone write this paper. A small nit is that I think the naming doesn't do the method justice. Why is this called "Longform"? Something nicer might be something like "reverse instructions" or prompt inversion.
-
Causal Language Models: Seq2Seq Architecture Without Cross-Attention
By
–
Causal LMs are seq2seq models just with a causal mask and shared encoder decoder with no cross attention.
-

LLM Evaluation: The Biggest Open Research Challenge
By
–
Imo the biggest open research area in LLMs now is how to exactly do evaluation correctly. It's tricky as pointed out below
-
MMLU and Bigbench: Key AI Evaluation Metrics at Google
By
–
oh yea i just saw that. not sure what happened. but in general back at G we mainly looked at MMLU and Bigbench as key metrics. that said, i also don't have any access to investigate anything now as im not at G anymore.
-
LangChain Explained: What Does This AI Tool Actually Do?
By
–
"I'm still not completely sure what LangChain does, and at this point I'm too afraid to ask" Same lol.
-
Twitter Algorithm Research Curation Quality Concerns
By
–
i've gotten so many "how do u keep up with research" type of questions over the years. The answer is simple. You don't, you just sign up for an account for twitter and let the algorithm do the work for you. if you don't see the paper on twitter, maybe it's for a good reason.