Please start your DALL-E engines, and report back. Is DALL-E really this heteronormative? A reader writes, “picture of two men in love getting married, it will usually put a wife next to each of them that they are in love with.”
MULTIMODAL AI
-
Photorealistic Food Generation: Midjourney vs DALLE-3 Hyperrealism
By
–
Def photorealistic for food — midjourney, dalle 3 seem to take a more hyperrealistic bent
-
FSD predicts actions over frames for world model development
By
–
Looks like end to end FSD likely models future actions rather than future frames. And by focusing on predicting something less noise and more signal, it builds a good world model for a physical world generally intelligent agent.
-

DALL-E versus Midjourney: Image Quality and Realism Comparison
By
–
ikr! looks edible too, om nom nom. in contrast midjourney, even with raw mode on looks more like staged imagery:
-
AI Tool Hides Messages in Images: Subliminal Advertising Concerns
By
–
Esta herramienta basada en inteligencia artificial es capaz de camuflar mensajes dentro de imágenes y puede ser utilizada para impactarnos subliminalmente, lo que puede ser un problema.
— Juan Merodio (@juanmerodio) 15 octobre 2023
Si te interesa el tema de la publicidad subliminal no te pierdas este vídeo completo en mi… pic.twitter.com/XJE67GmRQkEsta herramienta basada en inteligencia artificial es capaz de camuflar mensajes dentro de imágenes y puede ser utilizada para impactarnos subliminalmente, lo que puede ser un problema. Si te interesa el tema de la publicidad subliminal no te pierdas este vídeo completo en mi
-

New Tech Trinity Temple Conceptualized by ChatGPT and Dall-E
By
–
OpenAI ChatGPT connection to Dall-E conceptualizes a temple dedicated to "The New Tech Trinity". Three colossal statues dominate the scene. Quantum Tech, carved from deep blue stone, levitates particles around him. Bio Tech, made of green and gold materials, cradles a biotech
-
Single Video vs Multi-View: Environmental Scanning Approach Requirements
By
–
Dope! Does this approach need synchronized multi-view video, or can it work with a single video and an environmental scan?
-
OpenAI Releases GPT-V System Card Technical Documentation
By
–
Worth reading: https://
cdn.openai.com/papers/GPTV_Sy
stem_Card.pdf
… -
Advanced AI Learning Path: Image Generation to Transformer Models
By
–
5 Introduction to Image Generation https://
cloudskillsboost.google/course_templat
es/541
… 6 Encoder-Decoder Architecture https://
cloudskillsboost.google/course_templat
es/543
… 7 Attention Mechanism https://
cloudskillsboost.google/course_templat
es/537
… 8 Transformer Models and BERT Model https://
cloudskillsboost.google/course_templat
es/538
… 9 Create Image Captioning Models https://
cloudskillsboost.google/course_templat
es/542
… -
MiniGPT-v2: Advanced Vision-Language Model Capabilities
By
–
👌Amazing work – MiniGPT-v2.
— Dr. Debashis Dutta (@debashis_dutta) 14 octobre 2023
It can perform many complex vision-language and visual grounding tasks with simple interaction👈
****************
🔗Demo & Project: https://t.co/u8zWBMGTjs
Team ✨️@garvinchen2 @tikgiau @xiaoqian_shen @lix709 @zechunliu @PengchuanZ Raghuraman… https://t.co/25HaG6RhBWAmazing work – MiniGPT-v2.
It can perform many complex vision-language and visual grounding tasks with simple interaction ****************
Demo & Project: http://
minigpt-v2.github.io
Team @garvinchen2 @tikgiau @xiaoqian_shen @lix709 @zechunliu @PengchuanZ
Raghuraman