ViCo: Detail-Preserving Visual Condition for Personalized Text-to-Image Generation paper page: https://
huggingface.co/papers/2306.00
971
… present a plug-in method, named ViCo, for fast and lightweight personalized generation. Specifically, we propose an image attention module to condition the
@_akhaliq
-

ViCo: Detail-Preserving Visual Condition for Personalized Text-to-Image Generation
By
–
-

Inserting Any Person into Diffusion Models with Single Photo
By
–
Inserting Anybody in Diffusion Models via Celeb Basis paper page: https://
huggingface.co/papers/2306.00
926
… propose a new personalization method that allows for the seamless integration of a unique individual into the pre-trained diffusion model using just one facial photograph and only 1024 -

Wuerstchen: Efficient Text-to-Image Model Pretraining Technique
By
–
Wuerstchen: Efficient Pretraining of Text-to-Image Models paper page: https://
huggingface.co/papers/2306.00
637
… introduce Wuerstchen, a novel technique for text-to-image synthesis that unites competitive performance with unprecedented cost-effectiveness and ease of training on constrained -

SafeDiffuser: Safe Planning with Diffusion Probabilistic Models
By
–
SafeDiffuser: Safe Planning with Diffusion Probabilistic Models paper page: https://
huggingface.co/papers/2306.00
148
… propose a new method, called SafeDiffuser, to ensure diffusion probabilistic models satisfy specifications by using a class of control barrier functions. The key idea of our -

ByteFormer: Transformers Processing Raw File Bytes for Image Classification
By
–
Bytes Are All You Need: Transformers Operating Directly On File Bytes paper page: https://
huggingface.co/papers/2306.00
238
… ByteFormer, achieves an ImageNet Top-1 classification accuracy of 77.33% when training and testing directly on TIFF file bytes using a transformer backbone with -

Brainformers: Trading Simplicity for Efficiency in Transformers
By
–
Brainformers: Trading Simplicity for Efficiency paper page: https://
huggingface.co/papers/2306.00
008
… develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms -

Understanding Transformer Internal Mechanisms Through Memory Analysis
By
–
Birth of a Transformer: A Memory Viewpoint paper page: https://
huggingface.co/papers/2306.00
802
… Large language models based on transformers have achieved great empirical successes. However, as they are deployed more widely, there is a growing need to better understand their internal mechanisms -

Analyzing Attention Glitches in Transformer Language Models
By
–
Exposing Attention Glitches with Flip-Flop Language Modeling abs: https://
arxiv.org/abs/2306.00946 identifies and analyzes the phenomenon of attention glitches, in which the Transformer architecture's inductive biases intermittently fail to capture robust reasoning. To isolate the -

Hiera: Hierarchical Vision Transformer Simplified Architecture
By
–
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles abs: https://
arxiv.org/abs/2306.00989 Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to -
Understanding Concept Representations in Text-to-Image Diffusion Models
By
–
The Hidden Language of Diffusion Models
— AK (@_akhaliq) 2 juin 2023
paper page: https://t.co/biwX1KX7hG
tackle the challenge of understanding concept representations in text-to-image models by decomposing an input text prompt into a small set of interpretable elements. This is achieved by learning a… pic.twitter.com/mF0QzNxAjoThe Hidden Language of Diffusion Models paper page: https://
huggingface.co/papers/2306.00
966
… tackle the challenge of understanding concept representations in text-to-image models by decomposing an input text prompt into a small set of interpretable elements. This is achieved by learning a