A 3 Trillion (yes, trillion) token open-source training dataset for LLMs: Dolma is built from a mixture of web content, scientific papers, code, public-domain books, social media, and encyclopedic materials. Now available on Huggingface if you have a spare 5.4 Terabytes
DATA
-

AI2 Launches OLMo Open Language Model Framework
By
–
@allen_ai launched "OLMo: Open Language Model". AI2 opens its framework that provide access to data, training code, models, and evaluation code. The training data Dolma is also released, a dataset of 3 trillion tokens from a diverse mix of web content, academic
-
Ludwig Celebrates 10k Milestone with LLM Hackathon Winners
By
–
Congrats to our #LLM Hackathon Winners! To celebrate Ludwig's 10k we ran a mini virtual hackathon. Projects include: Automating data quality checks in Python #QA for Indian Tax Code #Classification for Customer Service
and more! https://
pbase.ai/3HGSQMo -

Building Long-Context RAG with Nomic Embeddings and Mistral
By
–
Build Long-context RAG from scratch: Nomic Embeddings + Mistral The context window of open source LLMs and embedding models has been relatively small vs proprietary models. But, methods to expand context window (RoPE, self-extend) are quickly changing this. Today, @nomic_ai has
-

Custom LLM and AI Agents for Structured and Unstructured Data
By
–
An AI Brain for your organization – custom LLM and AI Agents (RAG) on Structured + Unstructured Data Imagine having a ChatGPT-like interface over all your structured (database) and unstructured data. You can ask a question to an AI bot, and it can run multiple parallel queries,
-

Nomic Embed: Open Text Embedding Model and Dataset Released
By
–
if you're doing research on text embeddings, you know that there are lots of tricks required to train a good model, and no open datasets.
— dr. jack morris (@jxmnop) 1 février 2024
we trained Nomic Embed, a great text embedding model, and actually released the data! certainly will make my research easier – check it out! https://t.co/akhYHzjKbBif you're doing research on text embeddings, you know that there are lots of tricks required to train a good model, and no open datasets. we trained Nomic Embed, a great text embedding model, and actually released the data! certainly will make my research easier – check it out!
-
Sharing Python, Data Science, ML and LLM Content Daily
By
–
That's a wrap! Every day, I share and simply content around Python, Data Science, Machine Learning & Large Language Models. Find me → @Sumanth_077 Like/RT the first tweet and help this reach more people.
-
Mastercard Launches Proprietary AI Model for Fraud Detection
By
–
Exclusive: Mastercard has launched is own proprietary generative AI model it says can boost fraud detection by up to 300%. https://
cnbc.com/2024/02/01/mas
tercard-launches-gpt-like-ai-model-to-help-banks-detect-fraud.html
… -
Teknium releases high-quality AI data filtering and curation
By
–
great release from @teknium
, doing a lot of high-quality data filtering and curating work, and opening it up. -
H2O Enterprise GPTe: Secure Generative AI for Internal Data
By
–
Check out our latest YouTube short, from our head of global training Andreea Turco @GreetData
, introducing H2O Enterprise GPTe! Discover how it simplifies generative AI use, ensuring secure internal data access and informed decision-making:
