SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis paper page: https://
huggingface.co/papers/2307.01
952
… present SDXL, a latent diffusion model for text-to-image synthesis. Compared to previous versions of Stable Diffusion, SDXL leverages a three times larger UNet
MULTIMODAL AI
-

SDXL: Advanced Latent Diffusion Model for High-Resolution Image Synthesis
By
–
-

DiT-3D: Diffusion Transformers for 3D Shape Generation
By
–
DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation paper page: https://
huggingface.co/papers/2307.01
831
… Recent Diffusion Transformers (e.g., DiT) have demonstrated their powerful effectiveness in generating high-quality 2D images. However, it is still being determined -

Training GPT-4 Style Language Models with Multimodal Inputs
By
–
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs? paper page: https://
huggingface.co/papers/2307.02
469
… Recent advancements in Large Language Models (LLMs) such as GPT4 have displayed exceptional multi-modal capabilities in following open-ended instructions given -
Apple’s New Siri Still Falls Short Despite LLM Work
By
–
New Siri comes with Vision Pro. But it will still suck. Apple is working on LLMs now, but that will take a while to play out.
-

AI Training Data Uses Unauthorized Facial Recognition Images
By
–
a thing that's crazy to me is not that my image been used for training AI image datasets, but that images of images of me from @pimeyes
' facial-recognition software (which i allowed to be included in a story i wrote about FR for @CNN
) have been used for training. -

AI Video Generation Tools: Latest Advancements and Capabilities
By
–
Las ia que crean imágenes estáticas han mejorado mucho y nos ofrecen resultados altamente creíbles y muy impactantes, pero en el caso del vídeo las cosas van un poco más despacio. Hoy te comparto dos de las herramientas más avanzadas en este campo y que ya, a día de hoy, son
-

Vision System Combines Image Generation and Recognition via Masking
By
–
Vision system marries image generation and recognition by inferring the missing parts of an image. It uses a variable masking strategy during pre-training, helping a neural net fill in the gaps for tasks like editing and learning: https://
bit.ly/3r4QL7Y -
OpenFlamingo v2: Advanced Multimodal Models Released
By
–
Research blog post: OpenFlamingo v2: Next-level models and enhanced training setup! The paper by @anas_awadalla and @irena_gao
, with compute support from Stability AI, details five trained OpenFlamingo models across the 3B, 4B, and 9B scales. #StabilityAI Read more: -
Robots Storytelling: How AI Shapes Narratives
By
–
Entertain the robots and they will tell your story to others.