Example here is the llm.c GPT-3 (124M) training on FineWeb (figure cropped at 250B tokens), we seem to surpass GPT-3 HellaSwag (green line) at ~150B tokens, per paper expected this to be at 300B tokens. Will re-run with FineWeb-Edu. I do want to be a bit careful on conclusions
OPEN SOURCE
-
llm.c Outperforming GPT-2/3 with Fewer Training Tokens
By
–
In llm.c pretraining we were already mildly perplexed why seem to be outperforming GPT-2 & 3 (124M) training on just 10B tokens instead of something closer to 100-300B, per the original papers. I suspect a good chunk of it may be just the dataset quality, so I'm eager to retrain
-

FineWeb-Edu: High-Quality LLM Dataset Filtering for Better Learning
By
–
Awesome and highly useful: FineWeb-Edu High quality LLM dataset filtering the original 15 trillion FineWeb tokens to 1.3 trillion of the highest (educational) quality, as judged by a Llama 3 70B. +A highly detailed paper. Turns out that LLMs learn a lot better and faster
-

Awesome LLM Apps with RAG and AI Agents Repository
By
–
Thank you for sharing. Find all the awesome LLM apps with RAG and AI agents in this opensource repository.
-
Phidata Library Praised for AI Agent Development
By
–
Thank you for sharing. Phidata is super cool library
-

Microsoft Launches Free 18-Lesson Generative AI Course on Github
By
–
Microsoft launched the best course on Generative AI! The free 18 lesson course is available on Github and will teach you everything you need to know to start building Generative AI applications.
-
U-Net Semantic Segmentation Library Released
By
–
How it works Internally it uses a U-Net based foreground/background semantic segmentation and yields the post processed results. If you try it, library would download the trained model first. Fine the code here:
-

Python Library Removes Image Backgrounds in Five Lines
By
–
This Python library is so powerful! Remove Background of an Image in just 5 lines of code! Try yourself!
-
SimPO: Reference-Free Preference Optimization for Language Models
By
–
10/ SimPO – a simpler and more effective approach for preference optimization with a reference-free reward; uses the average log probability of a sequence as an implicit reward (i.e., no reference model required) which makes it more compute and memory efficient.
-

Solar-1-mini-chat-ja: Compact Japanese LLM Model
By
–
#solarllmja is registered at #awesome_japanese_llm, https://
github.com/llm-jp/awesome
-japanese-llm
…. If you have any service in Japan , please checkout `solar-1-mini-chat-ja` at https://
console.upstage.ai. It's small but very string in Japanese!
