(3/3) Our new Multimodal Diffusion Transformer (MMDiT) architecture uses separate sets of weights for image and language representations, which improves text understanding and spelling capabilities compared to previous versions of SD3. Read the full blog post and access the
Stable Diffusion 3 MMDiT Architecture Enhances Text Understanding
By
–
