Controlling Text-to-Image Diffusion by Orthogonal Finetuning paper page: https://
huggingface.co/papers/2306.07
280
… Large text-to-image diffusion models have impressive capabilities in generating photorealistic images from text prompts. How to effectively guide or control these powerful models to
AI
-

Controlling Text-to-Image Diffusion by Orthogonal Finetuning
By
–
-

Face0: Instant Face Conditioning for Text-to-Image Models
By
–
Face0: Instantaneously Conditioning a Text-to-Image Model on a Face paper page: https://
huggingface.co/papers/2306.06
638
… present Face0, a novel way to instantaneously condition a text-to-image generation model on a face, in sample time, without any optimization procedures such as fine-tuning or -
Humorous commitment to converting someone to PyTorch
By
–
maybe not this year, but I'll convert you to PyTorch eventually haha
-
Pretrained Models Make Fine-tuning the Only Practical Approach
By
–
Hah, and even further, with pretrained models available, anything more than finetuning feels like a pain
-
Schmidhuber’s 1991 Alternative to RNNs: The Origins of Linear Transformers
By
–
Hah, yeah, little known fun fact: Schmidhuber proposed an alternative to RNNs back in 1991, which is now called "linear Transformers" or "Transformers with linearized self-attention" via more recent papers. Summarized it here: https://
magazine.sebastianraschka.com/p/why-the-orig
inal-transformer-figure
… -

Transformer’s Birthday: Attention Mechanisms and RNN Evolution
By
–
Happy birthday, transformer! An awesome summary @DrJimFan
! Also interesting to think about why we needed attention for RNNs (before transformers) in the first place. Since we can't translate word-by-word, we needed a RNN encoder-decoder setup. But then, it's hard to remember. -

Augmenting Language Models with Long-Term Memory
By
–
Augmenting Language Models with Long-Term Memory paper page: https://
huggingface.co/papers/2306.07
174
… Existing large language models (LLMs) can only afford fix-sized inputs due to the input length limit, preventing them from utilizing rich long-context information from past inputs. To address -

Weakly Supervised Information Extraction from Handwritten Document Images
By
–
Weakly supervised information extraction from inscrutable handwritten document images paper page: https://
arxiv.org/abs/2306.06823 State-of-the-art information extraction methods are limited by OCR errors. They work well for printed text in form-like documents, but unstructured, -

High-Fidelity Audio Compression with Improved RVQGAN
By
–
High-Fidelity Audio Compression with Improved RVQGAN paper page: https://
huggingface.co/papers/2306.06
546
… Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can -

Aladdin: Zero-Shot 3D Asset Generation from Scene Descriptions
By
–
Aladdin: Zero-Shot Hallucination of Stylized 3D Assets from Abstract Scene Descriptions paper page: https://
huggingface.co/papers/2306.06
212
… What constitutes the "vibe" of a particular scene? What should one find in "a busy, dirty city street", "an idyllic countryside", or "a crime scene in an