quite incredible to see the goalposts of long context move in the past 1 year. May 2023: asking @jefrankle and @abhi_venigalla about their 65k+ "Llongboi" model May 2024: @markatgradient casually extending Llama 3 to >1m tokens with ~perfect NIAH and mainstream ai engineers
@swyx
-
Questioning the legitimacy of Alibaba virtual try-on models
By
–
oh you do virtual try ons! just wondering – have you seen if any of the Alibaba models are legit? or are they vaporware?
-

Jeremy Howard’s Third Latent Space Podcast Appearance Coming Soon
By
–
we just recorded @jeremyphoward
's 3rd appearance on @latentspacepod (after the NeurIPS guest spot https://
x.com/latentspacepod
/status/1741160693582504275
…) and i have to say this one was EVEN MORE BANGING than the last. the world is not ready for what jeremy is about to drop in the next few months -
Curated AI Podcasts and Newsletters Resource List
By
–
periodic reminder that i do maintain a list of technically focused ai podcasts and newsletters here https://
github.com/swyxio/ai-note
s/blob/main/Resources/GoodAIPodcastsandNewsletters.md
… obviously its hard to cover math and images on voice format but the people who #learninpublic are heroes for helping the rest of us keep up -

GPT-3 Hyperparameters Analysis and Model Scaling Expectations
By
–
ah ok these are gpt3 hparams. it sounds like this alone would be enough to beat GPT2, and its unknown if 290B more tokens would get it to match or beat GPT-3 Small looking forward to the 1.5b – fascinating to see these all documented and taught live!!
-
FineWeb Dataset: Improvements Over GPT-2 Training Data
By
–
10B tokens of FineWeb! Ilya said WebText was 40B tokens (
https://
youtube.com/watch?v=13CZPW
mke6A&t=3645s
… – for gpt2 1.5b) what accounts for the improved loss/accuracy that you got over GPT2 – have we improved our dataset filtering? were there smarter hparam choices made here? any ballpark attributions -
ICLR Episode Part 1 Released, Preparing for AI.Engineer Conference
By
–
ICLR episode part 1 just shipped! technically still in time for the long weekend haha
— swyx 🐣 (@swyx) 27 mai 2024
on a meta level it’s been great studying all the big ML conferences in prep for @aiDotEngineer next month. Lots of good ideas for running >5000 person affairs with something for everyone. Any… https://t.co/kiKBcJiQe6 pic.twitter.com/6h91b4D0dXICLR episode part 1 just shipped! technically still in time for the long weekend haha on a meta level it’s been great studying all the big ML conferences in prep for @aiDotEngineer next month. Lots of good ideas for running >5000 person affairs with something for everyone. Any
-

Sharing LLM OS Blog Post at AI Engineering Singapore Event
By
–
had a fun time mouthblogging my LLM OS blogpost I've been cooking on for the past month 🙂 AI Eng Singapore is alive and well thanks to Gabriel and @ivanleomk and all the other tech scene friends! https://t.co/rB5rPbXoKb pic.twitter.com/qDUhX6ayEB
— swyx 🐣 (@swyx) 25 mai 2024had a fun time mouthblogging my LLM OS blogpost I've been cooking on for the past month 🙂 AI Eng Singapore is alive and well thanks to Gabriel and @ivanleomk and all the other tech scene friends!
-

LLM Bugs: Finding and Discussing Detection Methods
By
–
LLM bugs and how to find them Part of the joy of curating @aidotengineer is I get to bring some of the best people in the community that you only see online, to meet and teach in person. Several of our speakers are giving their first talks -ever-! Daniel is one of the most
-
The Overwhelming Learning Curve for New AI Enthusiasts
By
–
honestly the eternal september effect in ai is going to be exhausting bc there is so much lore to learn. i dont envy people who are just coming onboard now. need introductory courses like mine and Hamel's to catch up
