somehow the corners of Twitter have made it about IITs and VC and something something. What you really need is to suffer with setting up the cuda version you *actually* want on Runpod.
CODE
-
Microsoft Expands Copilot Access to Developers Everywhere
By
–
#Copilots For Everyone: Microsoft Brings Copilots to the Masses! https://
devops.com/copilots-for-e
veryone-microsoft-brings-copilots-to-the-masses/
… via @devopsdotcom #DevOps #CloudNative #cloud #tech #lowcode #NoCode #DevSecOps #AIOps #MLOps #k8s #Kubernetes #Docker #github #Engineering #technology #OpenSource #bot #python #javascript -
SlimPajama Tools Released for Building AI Models
By
–
Here are the tools we created to build SlimPajama. Have fun! x.com/dmsobol/status…
-
SlimPajama Tools Released for Open Source AI
By
–
Here are the tools we created to build SlimPajama. Have fun!
-
Integrate with Hugging Face Transformers and Text Generation Inference
By
–
Nice! You should be integrate it with hf/transformers or https://
github.com/huggingface/te
xt-generation-inference
… -

AI Chat Platform for End-to-End ML and LLM Operations
By
–
Use our AI chat to program our end-to-end ML and LLM Ops platform – execute code – draw plots – analyze dataset – transform data – fine-tune LLMs You talk to the bot, and the AI bot builds ML models and deploys them in production.
-
Efficient LLM Fine-tuning with LoRA Tutorial on Keras
By
–
Awesome new tutorial on http://
keras.io: how to use LoRA to perform very efficient fine-tuning of LLMs. -
Custom Parallel Data Pipeline for Trillion Token Deduplication
By
–
It was no mean feat to deduplicate data on this scale – existing tools does not scale to a trillion tokens. We built a custom parallel data pre-processing pipeline and are sharing the code open source with the community.
-

SlimPajama: 50% Smaller, Twice Faster LLM Training Dataset
By
–
SlimPajama cleans and deduplicates RedPajama-1T, reducing the total token count and file size by 50%. It's half the size and trains twice as fast! It’s the highest quality dataset when training to 600B tokens and when upsampled performs equal or better than RedPajama.
-
SlimPajama: High-Quality Dataset Reduces Duplicates Training
By
–
RedPajama-1T is the largest open dataset today but contains a large percentage of duplicates, making a full training run costly and inefficient. Like the Falcon team, we found data quality is just as important as quantity – which led to SlimPajama.