Hello @sandbar101 thanks for your question. We are using https://
github.com/TencentARC/T2I
-Adapter
… which is very similar to controlnet, but works better.
MULTIMODAL AI
-
T2I-Adapter: Advanced Image Generation Technology Explained
By
–
-

Kandinsky 2.2 Model Now Available on Replicate Platform
By
–
Kandinsky 2.2 is live on Replicate. https://
replicate.com/cjwbw/kandinsk
y-2.2
… -
Stability AI Launches Stable Doodle Sketch-to-Image Tool
By
–
Exciting Announcement! Stability AI proudly presents Stable Doodle, our groundbreaking sketch-to-image tool!
— Stability AI (@StabilityAI) 13 juillet 2023
🖌️Unleash your creativity as a simple drawing transforms into a vibrant, dynamic image. #StabilityAI
Read more https://t.co/Oy4lGSv2e9 pic.twitter.com/vrGGlMY37fExciting Announcement! Stability AI proudly presents Stable Doodle, our groundbreaking sketch-to-image tool! Unleash your creativity as a simple drawing transforms into a vibrant, dynamic image. #StabilityAI Read more https://
bit.ly/44CV0Xr -

Robots Learning Manipulation from Human Video Demonstrations with Eye-in-Hand Cameras
By
–
Giving Robots a Hand: Learning Generalizable Manipulation with Eye-in-Hand Human Demonstrations paper page: https://
huggingface.co/papers/2307.05
959
… Eye-in-hand cameras have shown promise in enabling greater sample efficiency and generalization in vision-based robotic manipulation. -

NaViT: Vision Transformer for Flexible Aspect Ratios and Resolutions
By
–
Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution paper page: https://
huggingface.co/papers/2307.06
304
… The ubiquitous and demonstrably suboptimal choice of resizing images to a fixed resolution before processing them with computer vision models has not yet been -

SITTA: Semantic Image-Text Alignment for Image Captioning
By
–
SITTA: A Semantic Image-Text Alignment for Image Captioning paper page: https://
huggingface.co/papers/2307.05
591
… Textual and semantic comprehension of images is essential for generating proper captions. The comprehension requires detection of objects, modeling of relations between them, an -
Amazing team of scholars creating datasets and visualizing bird data
By
–
It's from an amazing team of scholars, coders & artists: -NYT & institutional data: @ananny -How to see birds: @JerThorp -C4 & the making of datasets: @_will_orr
-Unstable categories & biodiversity: @Hamsini_S
-The trouble with ImageNet: @SashaMTL -

Village of Whispers: Text to Video AI by Pika Labs
By
–
Village of Whispers, text to video AI, pikalabs by willis.visual pic.twitter.com/GRTNi6he8q
— AK (@_akhaliq) 12 juillet 2023Village of Whispers, text to video AI, pikalabs by willis.visual
-
ACL Outstanding Paper on Hybrid Transducer Attention Encoder-Decoder Speech
By
–
ACL Outstanding Paper
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks -

DeepFloyd IF: Advanced Text-to-Image Generation Model Released
By
–
Stable Diffusion models aren't our only text-to-image models. The @deepfloydai team worked hard to train and publicly release DeepFloyd IF, a state-of-the-art image generation model that is better at text and photorealism: