Welcome LeRobot! the first robotics library at Hugging Face Over the past months, we saw impressive research breakthroughs in robotics (ALOHA, diffusion policies, UMI and so many) enabling robots behaviors previously thought impossible to train with limited quantity of data
@thom_wolf
-

Hugging Face Making Waves in French Mainstream Media
By
–
I’m in “Paris Match” this week. With a big smile because things are going amazing at Hugging Face Paris Match is like the French equivalent of “Life” I guess – AI is really getting mainstream!
-
AI Company Claims Rare Sustainable Profitability Model
By
–
pretty much the opposite, we might be one of the (rare) sustainable companies in AI at the moment as we’re generally profitable
-

Budget 3D-Printed Teleoperation System Now Operational
By
–
A cheap 3d-printed teleoperation setup. Just finished it now and it’s working well. pic.twitter.com/zbB5Micfn7
— Thomas Wolf (@Thom_Wolf) 27 avril 2024A cheap 3d-printed teleoperation setup. Just finished it now and it’s working well.
-
Apple Releases OpenELM Open-Source Language Models Family
By
–
OpenELM: a family of Open-source Efficient Language Models Welcome Apple Inc. in the family of open-source LLM trainers! https://
huggingface.co/collections/ap
ple/openelm-instruct-models-6619ad295d7ae9f868b759ca
… And together with a new library: CoreNet https://
github.com/apple/corenet -
Simon Stålenhag’s Visionary AI and Technology Art
By
–
Haha came to say the things. Love the work of Simon Stålenhag
-
Open LLM Leaderboard Backstage: Insights and Development
By
–
A sneak peek in the backstage of the open LLM leaderboard by @clefourrier and @nathanhabib1011
-

Phi Models Outperform Larger Models: Open Textbook Initiative Needed
By
–
These new phi-1 & 1.5 models are fascinating, no? performances crushing 10 times bigger models secret sauce coming from a magic textbook dataset close to zero information on this dataset other than GPT3.5 generated Is it time for an "Open Textbook" project??
-
GPT-4 Falcon Llama2 Tokenizer Comparison Study Results
By
–
And our answers are out! Running on 1B tokens from the web (filtered and mostly in English as details in https://
huggingface.co/papers/2306.01
116
…) we got – GPT4 tokenizer (100k vocab) gives you 0.997B tokens – Falcon tokenizer (64k vocab) gives you ~5% more tokens (1.04B)
– Llama2 tokenizer -
Tokenizer vocabulary size comparison across language models
By
–
Sunday small guessing puzzle Let's say I have 3 tokenizers:
– llama2: 32k vocab
– falcon: 65k vocab
– GPT4: 100k vocab I take ~2M random documents from the web (let’s say 10 random parquet files from RefinedWeb from https://
huggingface.co/datasets/tiiua
e/falcon-refinedweb
… roughly 1B tokens). I tokenize them