New issue brief: Have Chinese AI models pulled ahead of their global counterparts? Our latest brief analyzes China’s diverse open-weight model ecosystem and examines the policy implications of their widespread global diffusion. https://
hai.stanford.edu/policy/beyond-
deepseek-chinas-diverse-open-weight-ai-ecosystem-and-its-policy-implications
…
GENERATIVE AI
-

Chinese AI Models Eclipse Global Competitors Policy Implications
By
–
-
Facebook’s new SAM audio model
By
–
And more 👀
— 🚨 AI News | TestingCatalog (@testingcatalog) 16 décembre 2025
Source https://t.co/hByoVPiVJF pic.twitter.com/jRHq39bcIHAnd more Source https://
about.fb.com/news/2025/12/o
ur-new-sam-audio-model-transforms-audio-editing/
… -
Meta’s SAM Audio Model for sound editing
By
–
Meta announced a new SAM Audio Model for audio editing that can isolate and extract any sound.
— 🚨 AI News | TestingCatalog (@testingcatalog) 16 décembre 2025
Quite clean 🔥 pic.twitter.com/2GtN2EAxYdMeta announced a new SAM Audio Model for audio editing that can isolate and extract any sound. Quite clean
-
Hindsight: Open-source solution for AI memory management
By
–
Ready to fix AI memory loss? Check out the repo and star the project here: https://
github.com/vectorize-io/h
indsight
… Full docs: -

ChatGPT Image 2 Model Rollout Begins
By
–
BREAKING : It looks like the Image 2 model rollout has started already for ChatGPT. It would also come along with a new UI for the Image tab with h style selector and a prompt bar. Did you get it too?
-
GPT and Claude struggle with self-improvement on visual tasks
By
–
Maybe gemini is better at this, but GPT and Claude models were pretty terrible at self-improving, if you give it a screenshot, its ability to actually pick up on obvious issues is pretty poor. I have tried this at the beginning, it came up with 10 different ways how to 'improve'
-

SWE-Playground: Synthetic Data Generation for Versatile Coding Agents
By
–

There are many good training methods for improving agents on SWE-bench: SWE-Gym, SWE-Smith, R2E-Gym. But what about broader software engineering tasks? In SWE-Playground, we introduce a new, more diverse synthetic data generation strategy to train divers software agents. Yiqi Zhu (@StephenZhu0218) Introducing SWE-Playground: A fully automated pipeline that generates synthetic environments to train versatile coding agents. 🤖✨ Training software engineering agents often relies on existing resources like GitHub issues and focuses on solving SWE-bench style issue resolution tasks. While this has driven incredible progress, real-world engineering involves a wider spectrum of tasks —from designing new libraries to writing reproduction scripts. 🌐 Rather than mining existing repositories, SWE-Playground synthetically generates projects, tasks, and verifiable unit tests from scratch. This approach offers two exciting opportunities: 1️⃣ Flexibility: We can generate tasks without being constrained by the availability or structure of existing open-source data. 2️⃣ Versatility: We extend training beyond Issue Resolution to include Issue Reproduction and Library Generation from Scratch. The results? 🚀 Our agents achieve strong performance across SWE-bench Verified, SWT-Bench, and Commit-0, demonstrating high data efficiency compared to baselines trained on larger datasets. Huge thanks to my amazing collaborators @apurvasgandhi and @gneubig for their incredible efforts on bringing this work to life! 👇 🧵 A deep dive into how we build versatile agents synthetically. Paper: arxiv.org/pdf/2512.12216 Project Page: neulab.github.io/SWE-Playgro… Code: github.com/neulab/SWE-Playgr… Data & Models: huggingface.co/collections/S… — https://nitter.net/StephenZhu0218/status/2000754124019683469#m
-

Express tutorial: working effectively with AI in 2026 using tools and prompts
By
–
How to work effectively with AI in 2026? In this express tutorial, I'll show you my complete method with my tools and my prompts → https://
youtu.be/CXNURuKWeJo -
![FLUX.2 [max] Released: Advanced Image Model with Web Search](https://artificialintelligencedynamics.com/wp-content/uploads/2026/05/xmon_5ba236d2_1777820165.jpg)
FLUX.2 [max] Released: Advanced Image Model with Web Search
By
–
FLUX.2 [max] is here. The highest performance image model from Black Forest Labs. Explore multi-image references, real-time context with web search, and high-fidelity image editing.
-

Frontier Models Struggle with Abstract Reasoning Tasks
By
–
NEW research on abstract reasoning. Frontier models like GPT-5 and Grok 4 still can't do what humans find trivially easy: infer transformation rules from a handful of examples. The default approach to solving ARC-AGI (the leading benchmark for abstract reasoning) treats these