Project: https://
github.com/deepseek-ai/Th
inking-with-Visual-Primitives
…
Paper: https://
github.com/deepseek-ai/Th
inking-with-Visual-Primitives/blob/main/Thinking_with_Visual_Primitives.pdf
…
MULTIMODAL AI
-

DeepSeek AI Releases Thinking with Visual Primitives Project and Paper
By
–
-

DeepSeek Releases Visual Primitives Reasoning Framework for Grounded Thinking
By
–
Huge! DeepSeek just released "Thinking with Visual Primitives" It's a reasoning framework that lets models “point” with visual markers (points, bounding boxes) while they think. Instead of describing locations in words, the AI grounds each step of its chain-of-thought directly
-
Vision Banana Unifies Object Detection Segmentation and Depth Estimation
By
–
What if image generators actually understand what they create?
— Satya Mallick (@LearnOpenCV) 30 avril 2026
Vision Banana proves it. One model handles:
→ Object detection
→ Instance segmentation
→ Metric depth from a single photo
→ Surface normal estimation
No specialist models. Beats SAM 3 on multiple segmentation… pic.twitter.com/mFY6lFTDC7What if image generators actually understand what they create? Vision Banana proves it. One model handles:
→ Object detection → Instance segmentation → Metric depth from a single photo → Surface normal estimation No specialist models. Beats SAM 3 on multiple segmentation -
Codex App Uses GPT-Image-2 to Design and Build Web Apps
By
–
You can just ask Codex to generate an image or a design in the Codex app! Our Build Web Apps plugin uses GPT-Image-2 automatically to design then build.
-

MacTok: AI Image Generation With Just 64 Tokens
By
–
Can AI generate images with just 64 tokens instead of thousands? Researchers from Fudan University introduce MacTok, a new continuous tokenizer. It uses clever image masking and representation alignment to prevent information loss, forcing the model to learn robust visuals
-
KAME Tandem Architecture Boosts Knowledge in Speech AI Systems
By
–
Two Heads Are Better Than One: Async Knowledge Injection for Speech AI with Tandem Architecture Blog: https://
pub.sakana.ai/kame/ KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI Paper: https://
arxiv.org/abs/2510.02327 #ICASSP2026 -

GPT Image 2 Viral Prompt Recreates Images in MS Paint Style
By
–

This GPT Image 2 prompt is going insanely viral right now. “Redraw the attached image in the most clumsy, scribbly, and utterly pathetic way possible. Use a white background, and make it look like it was drawn in MS Paint with a mouse. It should be vaguely similar but also not
-
OmniShotCut: Shot Boundary Detection with Transformer
By
–
OmniShotCut
— AK (@_akhaliq) 30 avril 2026
Holistic Relational Shot Boundary Detection with Shot-Query Transformer
paper: https://t.co/dNBF7bUeE8 pic.twitter.com/Q2qtWbeNVZOmniShotCut Holistic Relational Shot Boundary Detection with Shot-Query Transformer paper: https://
huggingface.co/papers/2604.24
762
… -
Agents and generative models solve blank canvas problem
By
–
It’s really awesome how agents and generative models have fully solved the blank canvas problem