[AINews 3 Apr 2026] Gemma 4: The world's best small Multimodal Open Models, dramatically better than Gemma 3 in every way https://
latent.space/p/ainews-gemma
-4-the-best-small-multimodal
… Congrats team!!
MULTIMODAL AI
-

Gemma 4: World’s Best Small Multimodal Open Models
By
–
-
OpenClaw Agents Now Support Video Calls via Google Meet
By
–
This is getting way too real!
— Shubham Saboo (@Saboo_Shubham_) 3 avril 2026
I can now get on a video call with my OpenClaw Agents to chat with them face to face.
All i need to do is to send them a Google meet invite. pic.twitter.com/3NnDA1p1nTThis is getting way too real! I can now get on a video call with my OpenClaw Agents to chat with them face to face. All i need to do is to send them a Google meet invite.
→ View original post on X — @saboo_shubham_, 2026-04-03 06:34 UTC
-
AI Video Calls Becoming Ordinary with Pika’s Real-Time Model
By
–
We’re just a couple years away from video calls with AI agents feeling completely ordinary. https://t.co/7USlSrH161
— Matt Shumer (@mattshumer_) 3 avril 2026We’re just a couple years away from video calls with AI agents feeling completely ordinary. Pika (@pika_labs) Conversations tend to go better with a face and a voice. That’s why we’re thrilled to release the beta version of the first video chat skill for ANY agent, powered by our new real-time model, PikaStream1.0. The skill preserves memory and personality, and enables real-time adaptability. And if you use it with your Pika AI Self, they’ll be able to execute agentic tasks during the call 💅 — https://nitter.net/pika_labs/status/2039804583862796345#m
→ View original post on X — @mattshumer_, 2026-04-03 04:31 UTC
-

Mastering PyTorch: Create and Deploy Deep Learning Models
By
–
"Mastering PyTorch: Create and deploy deep learning models from CNNs to multimodal models, LLMs, and beyond" – http://amzn.to/40IFEQR via @PacktDataML —————
#AI #ML #MachineLearning #DataScience #DataScientist #GenAI -

Streamo: Real-time Streaming Video LLM for Intelligent Assistance
By
–
What if an AI could truly understand live video streams and act as your intelligent assistant, in real-time? Researchers from Hong Kong Baptist University and Tencent Youtu Lab just unveiled a major step forward! They present Streamo, a real-time streaming video LLM. It's trained on a new, massive instruction dataset (Streamo-Instruct-465K) to enable unified understanding across many streaming video tasks. Streamo excels at real-time narration, complex action understanding, event captioning, and time-sensitive Q&A. It bridges the gap between static video analysis and genuinely interactive, intelligent multimodal AI assistants in continuous streams! Streaming Instruction Tuning Project: jiaerxia.github.io/Streamo/ Code: github.com/maifoundations/St… Our report: mp.weixin.qq.com/s/Q28azqwk-… 📬 #PapersAccepted by Jiqizhixin
→ View original post on X — @jiqizhixin, 2026-04-03 03:36 UTC
-

Build Text-to-Image Generator with Transformers and Diffusions
By
–
Build a Text-to-Image Generator (from Scratch), with transformers and diffusions: http://
amzn.to/3MFbyK4 by @mark_h_liu v/ @ManningBooks —————
#AI #ML #MachineLearning #DataScience #DataScientist -

Pika’s New Real-Time Video Chat Skill for AI Agents
By
–
Cool to see pika reinventing itself. Now I kinda wanna embody my open claw agent and jump into a real time video call. https://t.co/Oi98CNnjGl
— Bilawal Sidhu (@bilawalsidhu) 3 avril 2026Cool to see pika reinventing itself. Now I kinda wanna embody my open claw agent and jump into a real time video call. Pika (@pika_labs) Conversations tend to go better with a face and a voice. That’s why we’re thrilled to release the beta version of the first video chat skill for ANY agent, powered by our new real-time model, PikaStream1.0. The skill preserves memory and personality, and enables real-time adaptability. And if you use it with your Pika AI Self, they’ll be able to execute agentic tasks during the call 💅 — https://nitter.net/pika_labs/status/2039804583862796345#m
→ View original post on X — @bilawalsidhu, 2026-04-03 02:28 UTC
-
AI-Generated 24-Hour Content Channels: Grok Scripting to Imagine Production
By
–
I see a new kind of 24-hour-a-day channel. One on any topic. Your favorite sports team. The AI news of the day. The war in Iran. Keeping up with Elon. Etc etc etc. Let Grok really study lists in great detail. Write a script. Shove it over to Imagine. Build a show. Or a piece
-
Using smux to Orchestrate Claude Code and Codex in One Terminal
By
–
You can now make Claude Code and Codex talk with one terminal.
— AlphaSignal AI (@AlphaSignalAI) 3 avril 2026
smux is a tmux setup that lets Claude Code and Codex read, type, and trigger keys across panes.
That means two tools can pass work back and forth, reply inside the same workspace, and collaborate without APIs.
It… pic.twitter.com/IJMwp68eU8You can now make Claude Code and Codex talk with one terminal. smux is a tmux setup that lets Claude Code and Codex read, type, and trigger keys across panes. That means two tools can pass work back and forth, reply inside the same workspace, and collaborate without APIs. It
-
Running Gemma 4 with audio on Mac locally
By
–
Anyone figured out a recipe to run Gemma 4 E2B or E4B against audio files locally on a Mac yet? Omar Sanseviero (@osanseviero) The 2 small ones also support audio understanding! Including ASR, speech to translated text, and more — https://nitter.net/osanseviero/status/2039789969154183548#m