Finally, they added a wide diffusion head the DiTDH variant. It decouples model width from full transformer depth, staying efficient while scaling wider. Result: 2.16 FID on ImageNet-256. RAE-DiTDH outperforms every VAE-based diffusion model at every scale.
MULTIMODAL AI
-

DiT Convergence: Width Over Depth Matters
By
–
Diffusion Transformers struggled at first. Why? Their width was too small for RAE’s high-dimensional latents. The fix: scale width ≥ latent dimension. Once model width ≥ 768, DiT started converging instantly. Depth didn’t matter width unlocked training stability.
-

RAE + DINOv2-B beats VAEs with sharper images
By
–
Everyone assumed semantic encoders couldn’t reconstruct images.
Turns out, they can and better than VAEs. RAE + DINOv2-B achieves 0.49 rFID vs 0.62 for SD-VAE, with 6× less compute. That’s fewer FLOPs and sharper reconstructions. -

RAEs improve diffusion models over VAEs
By
–
Today, most diffusion models still use VAEs built on 2021 tech. They compress images into low-dimensional latents (like 4 channels). That’s why diffusion models lose global structure and texture fidelity. RAEs fix this by encoding rich semantic features directly from pretrained
-

New Paper Revolutionizes Diffusion Models
By
–
Holy shit…Diffusion just leveled up A new paper “Diffusion Transformers with Representation Autoencoders” basically kills the VAE era. Instead of the old VAE bottleneck, they use representation autoencoders (RAEs) built from pretrained encoders like DINO or SigLIP. The
-
Sora recreates iconic Gone with the Wind scene
By
–
Recreation of the Gone With the Wind scene by Sora. #Sora2 Follow me on Sora. (ID: kaifuleeai) pic.twitter.com/R13wvbK99o
— Kai-Fu Lee (@kaifulee) 18 octobre 2025Recreation of the Gone With the Wind scene by Sora. #Sora2 Follow me on Sora. (ID: kaifuleeai)
-
EgoAgent: First-Person AI Perception and Action Model
By
–
How can we build AI agents that perceive, predict, and act from a first-person view, just like humans?
— 机器之心 JIQIZHIXIN (@jiqizhixin) 18 octobre 2025
Inspired by the human perception-action loop, researchers propose EgoAgent, a unified agent model that simultaneously learns to represent the environment, predict the future,… pic.twitter.com/Gd4JWaZBTRHow can we build AI agents that perceive, predict, and act from a first-person view, just like humans? Inspired by the human perception-action loop, researchers propose EgoAgent, a unified agent model that simultaneously learns to represent the environment, predict the future,
-

PaDT: MLLMs Generate Visual Detection Outputs Directly
By
–
Ever wonder if an AI could do more than just describe an image and actually show you where things are? PaDT (Patch-as-Decodable Token) is a unified paradigm enabling Multimodal Large Language Models (MLLMs) to directly generate visual outputs like detection boxes and
-
Simulon builds advanced AI-powered image-based lighting estimation
By
–
I like light probes and I cannot lie. Especially when they’re being hallucinated by AI.
— Bilawal Sidhu (@bilawalsidhu) 17 octobre 2025
Simulon is building a much higher quality version of image-based lighting estimation the Vision Pro or ARCore does in real time. https://t.co/6lYCfFaxf7I like light probes and I cannot lie. Especially when they’re being hallucinated by AI. Simulon is building a much higher quality version of image-based lighting estimation the Vision Pro or ARCore does in real time.
-
Veo 3.1 Launch: New Creative Control Features and Nano Banana Integration
By
–
It’s been an unbelie-veo-ble week of launches! Here’s everything that happened: — We launched Veo 3.1 with a suite of new features to give you more creative control of your video generations. — We brought Nano Banana into new surfaces like @Google Search, @NotebookLM
, Slides in
