Meta team just dropped the first Watermarking model that not edit can break! Ever heard of watermarking? It's a technique that allows you to mark in an image its original source. It's our best shield against AI-generated deepfakes, or content stolen from artists!
@aymericroucher
-

Qwen2.5-Coder-32B: first open-source model to match GPT-4o
By
–
> Qwen2.5-Coder-32B: new best-in-class open coding model, beats GPT-4o on most coding benchmarks! It's the first time Open-Source coding model of this size class that clearly matches GPT-4o's coding capabilities! Completes the previous two Qwen 2.5 Coder release with 4
-

Scaling laws diminishing returns for OpenAI GPT models
By
–
> Are scaling laws over? A report from the Information announced that @OpenAI is seeing diminishing returns from scaling up the next GPT models. What are scaling laws? These are empiric laws that say "Every time you increase compute spent in training 10-fold, your LLM's
-
HfApiEngine improvement simplifies open-LLM agent creation
By
–
A nice PR from @BruleNaudet in transformers.agents just improved HfApiEngine. This makes the creation of open-LLM-powered agents even easier with our free Inference API!
-
Autogen-based agent team tops GAIA benchmark submissions
By
–
It's not a new framework, it's simply a team of agent based on Microsoft's autogen framework! And submissions of this structure from Microsoft have long been near the top of the GAIA benchmark.
-

AndroidLab benchmark shows small fine-tuned models can power JARVIS
By
–
> AndroidLab: First ever systematic benchmark for Android mobile agents shows that small, fine-tuned open models can power a JARVIS system on your smartphone A team from @Tsinghua_Uni just released AndroidLab, the first systematic framework to evaluate and train Android
-

Tencent releases Hunyuan-Large: open MoE model beats LLaMA 3.1-405B
By
–
> Hunyuan-Large just released by @TencentGlobal : Largest ever open MoE LLM, only 52B active parameters but beats LLaMA 3.1-405B on most academic benchmarks! Key insights: Mixture of Experts (MoE) architecture: 389 B parameters in total, but only 52B are activated for any
-
CLEAR: First Benchmark for Machine Unlearning
By
–
> CLEAR: first benchmark to make models forget what we want them to forget! With privacy concerns rising, we sometimes need our models to "forget" specific information – like a person's data – while keeping everything else intact. Researchers just released CLEAR, the first
-
Oasis: First AI-Generated Real-Time Video Game Without Engine
By
–
> Oasis: First Real-Time Video Game Without a Game Engine! 🎮@DecartAI & @Etched just released Oasis – a fully AI-generated video game running at 20 FPS (frames per second). The model takes keyboard inputs and generates everything – physics, rules, graphics – on the fly,… pic.twitter.com/DU6l3PBfhF
— m_ric (@AymericRoucher) 1 novembre 2024> Oasis: First Real-Time Game Without a Game Engine! @DecartAI & @Etched just released Oasis – a fully AI-generated video game running at 20 FPS (frames per second). The model takes keyboard inputs and generates everything – physics, rules, graphics – on the fly,
