Now given that next-word prediction is multi-task learning, we can write the overall loss is the weighted sum of loss of individual tasks. When overall loss improves smoothly, do all individual tasks improve smoothly, or do some improve at different rates than others?
LLMS
-

Scaling Laws: How Compute Investment Improves Model Performance
By
–
Next, “scaling laws” assert that overall loss is expected to improve as you scale up the compute used to train your model. This was the motivation for our current scaling paradigm—as you invest more in scale, your model reliably gets better instead of plateauing.
-

Predicting Next Word Reveals Arbitrary Tasks in Natural Text Distribution
By
–
Although the tasks above are well defined, the naturally occurring distribution of tasks in text turns out to be not well-defined. Here, I show that predicting the next word in a random sentence from Biden’s wikipedia page can include arbitrary tasks like comma prediction.
-

Language Models Learn Thousands of Tasks During Pre-training
By
–
The first step is to understand that when LMs are pre-trained on next-word prediction, they are really doing massive multi-task learning on thousands (millions?) of tasks. Here is a list of some potential tasks.
-
Whiteboard Lecture on Why Language Models Work Well
By
–
As a kid I loved whiteboard lectures way more than slides, so for Stanford’s CS25 class I gave a whiteboard lecture! My goal was to simply and clearly explain why language models work so well, purely via intuitions. Youtube video: https://
youtu.be/3gb-ZkVRemQ?si
=jvbzUmR9Q76PIc5r
… (w/ @hwchung27
) -
Yi Open Foundation Models Paper Receives Community Recognition
By
–
Thanks for highlighting our work, @rohanpaul_ai ! We're thrilled to see our "Yi: Open Foundation Models" paper resonating with the community.
-
Prompt Optimization: Haiku Outperforming Sonnet Model
By
–
Okay here's a fun one. Just optimized a prompt that consistently does better on Claude Haiku than Claude Sonnet even. There's so much about these LLMs we have yet to understand….
-
Daily Content on Python Data Science Machine Learning and MLOps
By
–
That's a wrap! If you are interested in any of these below topics: – Python – Data Science – Machine Learning – Data Analysis – LLMs – MLOps Find me → @Sumanth_077 I'm sharing daily content over here.
-
Mistral AI Raises 600 Million Euros in Funding
By
–
[#Article] @MistralAI confirms a fundraising round of 600 million euros https://actuia.com/actualite/mistral-ai-confirme-une-levee-de-fonds-de-600-millions-deuros/
… #AI #ArtificialIntelligence -

Google and OpenAI AI Model Release Timelines
By
–

DELAYED: New Gemini features will only be released next week on June 18. It looks like Google changed their forecast on when OpenAI will release something new
