GAIA: A benchmark for general AI assistants,
by a team from Meta-FAIR, Meta-GenAI, HuggingFace, and AutoGPT. Current Auto-Regressive LLMs don't do very well.
LLMS
-

GAIA Benchmark Tests Current Auto-Regressive LLM Performance
By
–
-
Karpathy’s Excellent LLM Guide Video Recommendation for Everyone
By
–
It's an excellent video by @karpathy that I just added to my guide for getting into LLMs. I recommend everyone listen to it regardless of their experience with AI. It will help you better understand the models and explain how LLMs work to non-tech people.
-
01.AI Launches Yi-34B-Chat Model with Quantized Versions
By
–
http://
01.AI team worked hard to launch Yi-34B-Chat finetuned on world's #1 open source base model, 4 bits & 8 bits quantized versions also went live on @huggingface
. More to build your LLM projects! -
OpenAI’s GPT5 Release Speculation in November 2023
By
–
It appears OAI already has enough innovation for GPT to release GPT5 lol. And November 2023 hasn’t even ended yet
-

Korean Open LLM Leaderboard Now Running on CPU Upgrade
By
–
Thanks to the generous support from @huggingface and the assistance of @clefourrier
, our Korean Open LLM board is now running on CPU UPGRADE! Now, LLMs on the Open Ko-LLM leaderboard will be evaluated at rocket speed! https://
huggingface.co/spaces/upstage
/open-ko-llm-leaderboard
… -
LightOn Unveils Alfred 2, Its Improved Open Source LLM Model
By
–
[#Article] LightOn announces the second version of #Alfred, its #LLM #opensource model Discover the improvements made to the second version https://actuia.com/actualite/lighton-annonce-la-seconde-version-dalfred-son-modele-llm-open-source/
… -
Training Costs of AI Models: OpenAI GTR and SBERT
By
–
yeah mostly bc they’re expensive to train, rn we have openAI GTR and sbert
-

New AutoTrain UI launches for easy LLM and model finetuning
By
–
A brand new AutoTrain UI just dropped! Now it's super duper easy to finetune LLMs, text classification models, image classification models, dreambooth lora, seq2seq and even tabular models! Powered by Hugging Face Spaces backend, the models take only a few seconds to load, so
-
LLM Evaluation Methods: Comparing Text Robustness Techniques
By
–
any reasonably competent LLM should be sufficient to evaluate, as LLMs are a lot more robust to comparing two texts (than to evaluate an answer in isolation). I'd say one can even just do cosine distance on sentence embeddings if we want to avoid the LLM evaluates an LLM
-
Training 2B-parameter AI systems to animal-level intelligence efficiently
By
–
About 2 billion neurons, like parrots, dogs, and octopus.
How do we get a machine with 2B neurons / 10T parameters to get as smart as octopus, dogs, parrots, and crows with just a few months' worth of real-time training data?