Gemma 4 also features multi-modal capabilities: use it with prompt + images combinations
MULTIMODAL AI
-
MultiGen: Level Design for Editable Multiplayer Worlds in Diffusion Engines
By
–
MultiGen
— AK (@_akhaliq) 3 avril 2026
Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines
paper: https://t.co/UT01uUYBxj pic.twitter.com/AMY0bDusAjMultiGen Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines paper: huggingface.co/papers/2603.0…
-

Visual Guide to Gemma 4: Exploring Google DeepMind’s New Models
By
–
A Visual Guide to Gemma 4 With almost 40 (!) custom visuals, explore the new models from Google DeepMind. We explore various techniques, ranging from Mixture of Experts and the Vision Encoder all the way up to Per-Layer Embeddings and the Audio Encoder. Link below 👇
→ View original post on X — @jeremyphoward, 2026-04-03 16:10 UTC
-
Try Lyria 3 Pro on Replicate
By
–
Try Lyria 3 Pro: replicate.com/google/lyria-3… [Translated from EN to English]
→ View original post on X — @replicate, 2026-04-03 15:59 UTC
-
Discover Lyria 3: Google’s New Audio Tool
By
–
Try Lyria 3:
replicate.com/google/lyria-3 [Translated from EN to English]→ View original post on X — @replicate, 2026-04-03 15:59 UTC
-
Seedance 2.0 Generates Stunning Kung Fu Video from Photos
By
–
Been messing with the new Seedance 2.0 on Higgs and it’s legitimately ace!
— Charly Wargnier (@DataChaz) 3 avril 2026
You feed it 2 photos and a prompt, and it outputs THIS 🤯
Nailing flawless kung fu physics + native audio from a couple of jpegs is just absurd https://t.co/9sod5mbZmh pic.twitter.com/CO0pQJ30BCBeen messing with the new Seedance 2.0 on Higgs and it’s legitimately ace! You feed it 2 photos and a prompt, and it outputs THIS Nailing flawless kung fu physics + native audio from a couple of jpegs is just absurd
-
Buzzy AI: Competing Agents Create Perfect Videos Automatically
By
–
With Buzzy, you don't babysit AI to create videos step by step.
— Buzzy Now (@Buzzy_now_AI) 3 avril 2026
You watch agents fight to compete for a perfect video.
Every battle teaches the system.
Every victory improves the next video.
Seedance 2 + Agent hunger game = Guaranteed Perfect Video pic.twitter.com/UMns08Tw8QWith Buzzy, you don't babysit AI to create videos step by step. You watch agents fight to compete for a perfect video. Every battle teaches the system. Every victory improves the next video. Seedance 2 + Agent hunger game = Guaranteed Perfect
→ View original post on X — @aihighlight, 2026-04-03 15:06 UTC
-

VLMgineer: AI-Powered Robots Design Their Own Tools
By
–
Can AI truly empower robots to invent their own solutions? George Jiayuan Gao, Tianyu Li, and colleagues from UPenn present VLMgineer. This framework leverages Vision Language Models (VLMs) to brainstorm initial tool designs and action plans. It then refines these ideas using evolutionary search in simulation, optimizing both the tool's geometry and how the robot uses it. VLMgineer consistently outperforms existing human-crafted tools and VLM-generated designs from human specifications across diverse, challenging everyday manipulation tasks, transforming complex robotics problems into straightforward executions. VLMgineer: Vision Language Models as Robotic Toolsmiths Project: vlmgineer.github.io Paper: arxiv.org/abs/2507.12644 Our report: mp.weixin.qq.com/s/FXdeQhAeq… 📬 #PapersAccepted by Jiqizhixin
→ View original post on X — @jiqizhixin, 2026-04-03 14:45 UTC
-
τ³-bench: Interactive Agent Evaluations for Knowledge and Voice
By
–
Really excited for the release of 𝜏³-bench, which brings interactive agent evals ever closer to real-world use cases across two dimensions: 1. 𝜏-knowledge evaluates agents that need to operate over noisy knowledge bases to figure out the correct policies/tools to use while serving a user 2. 𝜏-voice tests voice agents in interactive customer service style settings. If you are developing embedding or voice models for AI agents, 𝜏³ is a great testbed for you to see how your models would perform in a realistic downstream use case. Blog: sierra.ai/blog/bench-advanci… Tweets from @BenShi34 and @keshav_57: nitter.net/benshi34/status/203436… nitter.net/keshav_57/status/20346…
-

Microsoft’s MAI-Image-2 AI Model Generates Creative Images
By
–



Here are few image create using Microsoft's new ai image model – MAI-Image-2 with prompts.
