Nice! You should be integrate it with hf/transformers or https://
github.com/huggingface/te
xt-generation-inference
…
OPEN SOURCE
-
Integrate with Hugging Face Transformers and Text Generation Inference
By
–
-
Open science and open-source driving rapid AI progress
By
–
@Jason the reason it's going so fast is not because of AI using AI but because of open science and open-source IMO. PSA that if we stop open science and open-source it's going to slow down.
-
Building Quivr: AI-Powered Document Question Answering System
By
–
"Have fun and build" – @_StanGirard An awesome deep-dive on what it takes to build @quivr_brain – a service that can answer questions about any documents you upload Some fun viewing for a Friday afternoon 🙂
-
Efficient LLM Fine-tuning with LoRA Tutorial on Keras
By
–
Awesome new tutorial on http://
keras.io: how to use LoRA to perform very efficient fine-tuning of LLMs. -
RedPajama Project Launches with OpenTensor and Together Support
By
–
We’d like to thank our partner @opentensor for supporting this project. And credit goes to @togethercompute and the entire team that created the RedPajama dataset! We can’t wait to see what you’ll build. Join our Discord and let us know your feedback: https://
discord.com/channels/10859
60591052644463/1085960592050896937
… -
Custom Parallel Data Pipeline for Trillion Token Deduplication
By
–
It was no mean feat to deduplicate data on this scale – existing tools does not scale to a trillion tokens. We built a custom parallel data pre-processing pipeline and are sharing the code open source with the community.
-

SlimPajama: 50% Smaller, Twice Faster LLM Training Dataset
By
–
SlimPajama cleans and deduplicates RedPajama-1T, reducing the total token count and file size by 50%. It's half the size and trains twice as fast! It’s the highest quality dataset when training to 600B tokens and when upsampled performs equal or better than RedPajama.
-
SlimPajama: High-Quality Dataset Reduces Duplicates Training
By
–
RedPajama-1T is the largest open dataset today but contains a large percentage of duplicates, making a full training run costly and inefficient. Like the Falcon team, we found data quality is just as important as quantity – which led to SlimPajama.
-

SlimPajama-627B: Largest Deduplicated Open-Source LLM Dataset
By
–
New dataset drop!
Introducing SlimPajama-627B: the largest extensively deduplicated, multi-corpora, open-source dataset for training large language models. https://
cerebras.net/blog/slimpajam
a-a-627b-token-cleaned-and-deduplicated-version-of-redpajama
… -
Open Source AI Project Gains Real-World Adoption
By
–
wow that's very cool! love to see our open source used like this