Integrate it, test it, or benchmark it — this model opens the door for more accessible, high-performing multimodal AI. Explore it here:
MULTIMODAL AI
-

Baidu Open-Sources ERNIE-4.5-VL-28B Multimodal Model
By
–
Big news from @Baidu_Inc
! ERNIE-4.5-VL-28B-A3B-Thinking — a breakthrough lightweight multimodal reasoning model — is now officially open-sourced. It’s already trending #1 on Hugging Face’s Image-Text-to-Text models list. -
Breakthrough in Multimodal AI: Visual Analysis Efficiency
By
–
.visuals step by step — much like a human would. It can study a storyboard, analyze design drafts, or pull key insights from visual briefs, turning images into structured understanding.
This combination of efficiency + depth is what makes it a true breakthrough in multimodal AI. -
3B Parameter Model Matches GPT-5 and Gemini Performance
By
–
With just 3B active parameters, it matches Gemini-2.5-Pro and GPT-5-High, and even surpasses them on ChartQA and DocVQAval What can it achieve with only 3B active parameters?
For example, its “Thinking with Images” feature lets it zoom in, analyze details, and reason through -

Qwen Image Edit Plus LoRA Explorer Now Available
By
–
Qwen Image Edit Plus (2509) LoRA Explorer is here!
— Replicate (@replicate) 12 novembre 2025
Use any HF adapter fast.
Lightning fast.
Thanks to the @PrunaAI + @Alibaba_Qwen teams for making all of this possible pic.twitter.com/Vq99FWalzsQwen Image Edit Plus (2509) LoRA Explorer is here!
Use any HF adapter fast.
Lightning fast. Thanks to the @PrunaAI + @Alibaba_Qwen teams for making all of this possible -
Tavus Preview: 5 AI Agents for Work
By
–
BREAKING 🚨: Early preview of Tavus 👀
— 🚨 AI News | TestingCatalog (@testingcatalog) 12 novembre 2025
There, you will be able to work with 5 human-like agents via text, voice and video chats. Work apps can be connected as well, so AI agents can execute tasks autonomously.
The AI office 🤖 https://t.co/lKLTsl8pdE pic.twitter.com/sb4LgWxuxYBREAKING : Early preview of Tavus There, you will be able to work with 5 human-like agents via text, voice and video chats. Work apps can be connected as well, so AI agents can execute tasks autonomously. The AI office
-
ElevenLabs Releases Scribe V2 Realtime Transcription Model
By
–
ElevenLabs released Scribe V2 Realtime, a transcription low-latency model with support for 90+ languages. https://t.co/ksVHzto2uI pic.twitter.com/r6QJ1cWTTC
— 🚨 AI News | TestingCatalog (@testingcatalog) 11 novembre 2025ElevenLabs released Scribe V2 Realtime, a transcription low-latency model with support for 90+ languages.
-

Gemini Dynamic View introduces interactive canvas experiences
By
–

BREAKING : Early preview of Dynamic View on Gemini. Dynamic View (Creative Canvas) will let users render interactive canvas experiences inline within a Gemini Chat. What it can do:
– Search the web
– Generate Images
– Render Games and more! Magic box -

Templates on Sora, Veo, Aurora auto-generate product variants with hooks and CTAs
By
–
Each template runs on Sora 2, Veo 3.1, or Aurora depending on what you need. Cinematic product shots → Sora/Veo
UGC testimonials → Aurora You change your product once and it auto-generates variants with different hooks, CTAs, and platform ratios. -

Black Forest Labs preparing to release FLUX.2 AI model
By
–

BREAKING : Black Forest Labs is preparing to release FLUX.2 [pro] soon. It is expected to be available on both Playground and the API. Major upgrade