Started a collection with all the benchmark spaces I know on the hub and some description. Do you know other benchmark on the Hugging Face hub? ping me and I'll add them! https://
huggingface.co/collections/op
en-llm-leaderboard/the-big-benchmarks-collection-64faca6335a7fc7d4ffe974a
…
@thom_wolf
-

Big Benchmarks Collection: Comprehensive LLM Evaluation Resources Hub
By
–
-

Hugging Face Leaders Recognized in Time’s 100 AI List
By
–
Mind-blowed to see both my amazing co-worker @mmitchell_ai and awesome co-founder @ClementDelangue selected in Time's 100 AI Super proud of what @huggingface is creating for the community (PS lot of great people missing from the list as well)
-
Dr. Almazrouei Team Releases Amazing AI Project Update
By
–
Amazing release Dr. Ebtesam Almazrouei and the team
-
Multilingual AI Dataset Benchmark on HuggingFace Platform
By
–
That’s a really cool datasets and benchmark! Super nice to push progress on multilinguality. Would you be maybe interested in hosting there http://
HuggingFace.co/facebook/ and/or think about having an automatically updated community leaderboard around it -
34B Parameter Models Now Run on Laptops, Massive Hardware Progress
By
–
Crazy how 34 billions parameters models seemed huge and unmanageable outside of a data center just maybe 1.5 years ago. Now it’s laptop stuff https://t.co/gXCg770BOM
— Thomas Wolf (@Thom_Wolf) 31 août 2023Crazy how 34 billions parameters models seemed huge and unmanageable outside of a data center just maybe 1.5 years ago. Now it’s laptop stuff
-

WizardCoder vs Phind-V2: Code Model Performance Comparison
By
–
The WizardCoder vs Phind-V2 perf challenge is fascinating to watch! Both seems really good code models: https://
reddit.com/r/LocalLLaMA/c
omments/165qeb3/wizardcoder_vs_phindv2_prelim_pass1_comparison
… Funny to remember that GPT4 was only scoring 67% on HumanEval at its release OpenAI also went a long way in a few months up to its current 84%! -
Current LLM Hallucination Benchmarks: Evaluating Factual Accuracy
By
–
What interesting benchmarks on LLM hallucinations exist currently?
-
GPU alone doesn’t determine AI success against big tech
By
–
Strong take! I like how you challenge the open-source, that’s how it gets to do better! Not sure I follow you on the “GPU/TPU number is all that matter” though, if it was all, @midjourney or @runwayml would have never had any chance vs Google/OpenAI. But it’s not what we’ve seen
-
Hugging Face Learning Resources: Start Your AI Journey
By
–
The best places to start learning about the Hugging Face ecosystem are our courses: https://
huggingface.co/learn and the tasks page: https://
huggingface.co/tasks -

HumanEval Benchmark: Assessing AI Code Generation Capabilities
By
–
Time to read again Loubna’s nice post from last week diving in the HumanEval benchmark