Thanks so much for this award (makes me feel old )! I am honored to be sharing this award with amazing coauthors Decaf was the first open source version of AlexNet, and tested whether features learned by this amazing ImageNet classifier could be reused broadly in other
@oriolvinyalsml
-
Ahmad Al Dahle Team Launches Impressive New AI Model
By
–
Congrats @Ahmad_Al_Dahle and team, looks like a great model! @ylecun
, that's a fairly large off ramp -
Congratulations Karpathy on Your Announcement
By
–
Excellent news, congrats @karpathy
! Let us know if you need a pretty great LLM -

GraphCast Team Wins Award for Weather Prediction AI
By
–
Extremely proud of my partner in crime @FortunatoMeire and all of #GraphCast team for this award! The thing is heavy
-
Gemma 2 Launch: Compact Open Model Challenges Industry Giants
By
–
Gemma 2 has arrived! The arena for open models is also heating up. Extra exciting as 27B is ~3x smaller than Llama3 70B and ~10x smaller than NVIDIA's Nemotron 340B 🔥♊️💙 https://t.co/wk3drPb85O pic.twitter.com/78rQE8naaZ
— Oriol Vinyals (@OriolVinyalsML) 27 juin 2024Gemma 2 has arrived! The arena for open models is also heating up. Extra exciting as 27B is ~3x smaller than Llama3 70B and ~10x smaller than NVIDIA's Nemotron 340B
-

Model Evaluation Release: Importance of Held Out Testing
By
–
Evals are so important — congrats on the release! And good to see a truly "held out" test of our models. Sending some good Vibes your way : )
-
Code Size Inclusion in Language Model Computation
By
–
Size of the code must be included in overall LM size computation
-

Gemini 1.5 Pro Enters LMSys Arena as Top-Tier Model
By
–
Gemini 1.5 Pro has entered the (LMSys) Arena! Some highlights: -The only "mid" tier model at the highest level alongside "top" tier models from OpenAI and Anthropic -The model excels at multimodal, and long context (not measured here) -This model is also state-of-the-art
-
Data contamination detection script for training datasets
By
–
Just provide a script that will produce a number from 0 to 1 to be run on your training data on how contaminated it is. Everyone runs it and reports it as the first column of results.
-
HumanEval Benchmark Reveals Training Set Leakage Issues
By
–
HumanEval has become a benchmark that measures training set leakage. Which is actually quite useful knowledge
