Today, most diffusion models still use VAEs built on 2021 tech. They compress images into low-dimensional latents (like 4 channels). That’s why diffusion models lose global structure and texture fidelity. RAEs fix this by encoding rich semantic features directly from pretrained
GENERATIVE AI
-

New Paper Revolutionizes Diffusion Models
By
–
Holy shit…Diffusion just leveled up A new paper “Diffusion Transformers with Representation Autoencoders” basically kills the VAE era. Instead of the old VAE bottleneck, they use representation autoencoders (RAEs) built from pretrained encoders like DINO or SigLIP. The
-
Top 5 Technology Trends Reshaping 2026 Innovation
By
–
The Top 5 Technology Trends for 2026 From intelligent agents and space computing to bio-tech and digital twins, these are the five technological shifts set to reshape business, society and innovation in 2026. Read more https://
bernardmarr.com/the-top-5-tech
nology-trends-for-2026/
… #TechTrends #Innovation -
Claude Code coming to mobile app, hidden from public, almost ready
By
–
Vibe coders will soon be able to use Claude Code on the go, directly in the Claude mobile app!
— 🚨 AI News | TestingCatalog (@testingcatalog) 18 octobre 2025
It is almost ready for release and already working, but hidden from the public. pic.twitter.com/0DyHlVJXNFVibe coders will soon be able to use Claude Code on the go, directly in the Claude mobile app! It is almost ready for release and already working, but hidden from the public.
-
Sora recreates iconic Gone with the Wind scene
By
–
Recreation of the Gone With the Wind scene by Sora. #Sora2 Follow me on Sora. (ID: kaifuleeai) pic.twitter.com/R13wvbK99o
— Kai-Fu Lee (@kaifulee) 18 octobre 2025Recreation of the Gone With the Wind scene by Sora. #Sora2 Follow me on Sora. (ID: kaifuleeai)
-
alphaXiv Resources Tab Now Available for Research Papers
By
–
Check out http://
alphaXiv.org and click on the resources tab for any paper! -

GPT-6 will not arrive before the end of the year
By
–
No, GPT-6 is not coming before the end of the year. It is over, @patience_cave won
-

PaDT: MLLMs Generate Visual Detection Outputs Directly
By
–
Ever wonder if an AI could do more than just describe an image and actually show you where things are? PaDT (Patch-as-Decodable Token) is a unified paradigm enabling Multimodal Large Language Models (MLLMs) to directly generate visual outputs like detection boxes and
-
GPT-5 Solves Previously Open Mathematical Problems
By
–
Hi, as the owner/maintainer of http://
erdosproblems.com, this is a dramatic misrepresentation. GPT-5 found references, which solved these problems, that I personally was unaware of. The 'open' status only means I personally am unaware of a paper which solves it.
