Catch me up: How to try the new #AI #tech everyone is talking about https://
wapo.st/3IqLz4v via @washingtonpost
AI
-

How to Try the New AI Technology Everyone Is Discussing
By
–
-
RedPajama Project Launches with OpenTensor and Together Support
By
–
We’d like to thank our partner @opentensor for supporting this project. And credit goes to @togethercompute and the entire team that created the RedPajama dataset! We can’t wait to see what you’ll build. Join our Discord and let us know your feedback: https://
discord.com/channels/10859
60591052644463/1085960592050896937
… -
Custom Parallel Data Pipeline for Trillion Token Deduplication
By
–
It was no mean feat to deduplicate data on this scale – existing tools does not scale to a trillion tokens. We built a custom parallel data pre-processing pipeline and are sharing the code open source with the community.
-

SlimPajama: 50% Smaller, Twice Faster LLM Training Dataset
By
–
SlimPajama cleans and deduplicates RedPajama-1T, reducing the total token count and file size by 50%. It's half the size and trains twice as fast! It’s the highest quality dataset when training to 600B tokens and when upsampled performs equal or better than RedPajama.
-
Ace Pushes AI Art Generation to Next Level
By
–
Ace is taking AI Art to the next level. https://
x.com/aceiverse/stat
/aceiverse/status/1667242949925834752
… -
SlimPajama: High-Quality Dataset Reduces Duplicates Training
By
–
RedPajama-1T is the largest open dataset today but contains a large percentage of duplicates, making a full training run costly and inefficient. Like the Falcon team, we found data quality is just as important as quantity – which led to SlimPajama.
-
Perplexity Launches Personalized AI Profiles for Tailored Answers
By
–
Introducing your AI Profile on Perplexity! Personalize your AI answers with your bio, language, location, you name it. With dynamically generated questions, we're aiming to make your AI experience truly tailored to you. Dive in and experience Perplexity in a new personalized way… pic.twitter.com/HbmvJKMpZI
— Perplexity (@perplexity_ai) 9 juin 2023Introducing your AI Profile on Perplexity! Personalize your AI answers with your bio, language, location, you name it. With dynamically generated questions, we're aiming to make your AI experience truly tailored to you. Dive in and experience Perplexity in a new personalized way
-

SlimPajama-627B: Largest Deduplicated Open-Source LLM Dataset
By
–
New dataset drop!
Introducing SlimPajama-627B: the largest extensively deduplicated, multi-corpora, open-source dataset for training large language models. https://
cerebras.net/blog/slimpajam
a-a-627b-token-cleaned-and-deduplicated-version-of-redpajama
… -

Building a Tree-Structured Parzen Estimator from Scratch
By
–
Building a Tree-Structured Parzen Estimator from Scratch (Kind Of) https://
bit.ly/3OnbsWp #AI #MachineLearning #DeepLearning #LLMs #DataScience -
Realistic AI Views: Opportunity Between Hype and Skepticism
By
–
Anyhow — there is still a world of opportunity ahead for those with realistic views about the potential of AI. Believing in magic and fairies is in every way as counter productive as being a tech reactionary.