This dataset is a very famous NLP summarization benchmark which has a long history being originally created by DeepMind researchers, adapted by MILA and IBM and in its latest version by a PhD student from Stanford. We host the same access to it as the one in the tensorflow
@thom_wolf
-
High Quality Corpus Training Efficiency for Language Models
By
–
There seems to be some growing hints that a few epochs of training on a high quality (and well deduplicated corpus) is just fine
-

Training Smaller Models Longer Challenges Chinchilla Predictions
By
–
There is a fascinating recent trend of training *smaller models for longer* w.r.t. Chinchilla optimal predictions Best explanation I've seen of this? This new blog post by @harm_devries (with collaborators of the @BigCodeProject
): https://
harmdevries.com/post/model-siz
e-vs-compute-overhead/
… Clearly these are only -

Funny AI Model with RLHF Tutorial on Hub
By
–
possibly one of the funniest model on the hub (the RLHF tutorial that was the reason of its creation is also awesome – check it out)
-
First Open-Source AI Riot Sparked 2023 Revolution
By
–
[Narrator] "…and the first open-source riot that lead to the revolution we all know started in spring 2023, from what came to be known as the largest AI meetup ever…"
-

GPT-4 Sparks of AGI Study Reveals Mind-Blowing Examples
By
–
There are completely mind-blowing examples in the GPT4 "sparks of AGI" study
-
Tech Company Stops Knowledge Sharing Platform Hosting Model
By
–
Not really since we are a hosting/sharing platform. I’m more a bit annoyed that they’ve stopped sharing knowledge but it’s fair game in their position I guess
-
Technical Reports Becoming Marketing Press Releases
By
–
when your technical report is actually more of a product marketing press release
-
Building Community Knowledge Over Isolated AI Silos
By
–
would be amazing to continue walking this road as a community and not isolated knowledge islands
-
Fake It Until You Make It: The AGI Development Philosophy
By
–
Fake it (AGI) until you make it (AGI)