Training t5 on lm objective for 100k more steps. There are lm adapted ckpts of t5 on the internet iirc.
CODE
-
Tokenizer Effects on Math Tasks: T5 vs LM Adaptation
By
–
Math tasks could be due to c4 sentencepiece/tokenizer though. T5 can't open end generate that well due to span corruption objective but LM adaptation or ul2 fixes this
-
SlimPajama Dataset: Preprocessing Library for LLM Training
By
–
SlimPajama dataset – https://
lnkd.in/gCchZ-xz
Preprocessing library: https://
lnkd.in/gV7r3YNC
Read our blog: -

SlimPajama: Open-Source Cleaned RedPajama Dataset Released
By
–
We recently announced the availablity of SlimPajama – an open-source, cleaned, and deduplicated version of RedPajama-1T. It is half the size and trains twice as fast and when upsampled, performs equal or better than RedPajama. See below for the dataset and preprocessing library
-

askgpt: AI-Powered Chat Interface for Learning R Programming
By
–
Introducing `askgpt`: a chat interface that helps you to learn R! | Johannes B. Gruber https://
bit.ly/3Mp9GBc #AI #MachineLearning #DeepLearning #LLMs #DataScience -
Programming with GPT Simplifies Developer Workflow
By
–
Programming with GPT just makes my life a lot easier
-
UK government gets access to major AI models
By
–
The UK government was just granted access to the source code for all major AI models: "Today, Google DeepMind, OpenAI and Anthropic have agreed to open up their AI models to the U.K. government – for research and safety purposes."
-

Deep Learning Lecture: From RNNs to Transformer Architecture
By
–
Link to the lecture in case someone's curious: https://
lightning.ai/pages/courses/
deep-learning-fundamentals/unit-8.0-natural-language-processing-and-large-language-models/8.4-from-rnns-to-the-transformer-architecture/
… -
Comparison of BoW and BERT Models on Same Dataset
By
–
Yes, the two links above. It's both on the same dataset. One is a BoW model, one is a BERT model.
-
Chat Completion API compatibility question
By
–
super cool!! does it work with the Chat Completion API?