I'll be in Davos next week as part of the Unicorn program at the World Economic Forum Let's meet if you're interested in leveraging/building/discussing AI!
@thom_wolf
-

Agentic Evals: Unlocking Hidden Capabilities of AI Models
By
–
when people ask whether you need an agent framework at all – mind blown all evals should move to agentic evals in 2025 imo we’re just leaving so much capabilities of our models on the table…
-
Open Source Vision Language Models: SmolVLM, MOLMO, Qwen VL
By
–
You should use a VLM. SmolVLM, MOLMO or Qwen VL are good open source starting points little bot
-

SmolLM2 Evaluation: Not Tuned for LMSys Arena
By
–
and interestingly here, SmolLM2 is not specifically tuned for LMSys Arena like most new models these days
-

LLM Evals Market Saturation Before Academic Publication
By
–
new LLM evals entering and leaving the field saturated before the paper is even published in ML conferences pic.twitter.com/S6zG9JUD1P
— Thomas Wolf (@Thom_Wolf) 6 janvier 2025new LLM evals entering and leaving the field saturated before the paper is even published in ML conferences
-
Most impactful AI evaluation releases in 2024
By
–
What was the most impactful/visible/useful release on evaluation in AI in 2024?
-

FineWeb Dataset Updated With December 2024 Knowledge Cutoff
By
–
Big update for teams pretraining LLMs: the FineWeb dataset has just been updated with a knowledge cut-off at Dec 31th 2024
-

Hugging Face Releases smolagents Open Source Library for AI Agents
By
–
When you release a new lib on December 31 and the community gets crazy Check it out if you're interested in agents: smolagents – a barebones library for agents – agents write python code to call tools and orchestrate other agents. Github: https://
github.com/huggingface/sm
olagents
… And the
