Retrieval-Enhanced Contrastive Vision-Text Models paper page: https://
huggingface.co/papers/2306.07
196
… Contrastive image-text models such as CLIP form the building blocks of many state-of-the-art systems. While they excel at recognizing common generic concepts, they still struggle on
@_akhaliq
-

Retrieval-Enhanced Contrastive Vision-Text Models for Improved Concept Recognition
By
–
-

FasterViT: Fast Vision Transformers with Hierarchical Attention
By
–
FasterViT: Fast Vision Transformers with Hierarchical Attention paper page: https://
huggingface.co/papers/2306.06
189
… design a new family of hybrid CNN-ViT neural networks, named FasterViT, with a focus on high image throughput for computer vision (CV) applications. FasterViT combines the -

Benchmarking Neural Network Training Algorithms: Performance Analysis
By
–
Benchmarking Neural Network Training Algorithms paper page: https://
huggingface.co/papers/2306.07
179
… Training algorithms, broadly construed, are an essential part of every deep learning pipeline. Training algorithm improvements that speed up training across a wide variety of workloads (e.g., -

Controlling Text-to-Image Diffusion by Orthogonal Finetuning
By
–
Controlling Text-to-Image Diffusion by Orthogonal Finetuning paper page: https://
huggingface.co/papers/2306.07
280
… Large text-to-image diffusion models have impressive capabilities in generating photorealistic images from text prompts. How to effectively guide or control these powerful models to -

Face0: Instant Face Conditioning for Text-to-Image Models
By
–
Face0: Instantaneously Conditioning a Text-to-Image Model on a Face paper page: https://
huggingface.co/papers/2306.06
638
… present Face0, a novel way to instantaneously condition a text-to-image generation model on a face, in sample time, without any optimization procedures such as fine-tuning or -

Augmenting Language Models with Long-Term Memory
By
–
Augmenting Language Models with Long-Term Memory paper page: https://
huggingface.co/papers/2306.07
174
… Existing large language models (LLMs) can only afford fix-sized inputs due to the input length limit, preventing them from utilizing rich long-context information from past inputs. To address -

Weakly Supervised Information Extraction from Handwritten Document Images
By
–
Weakly supervised information extraction from inscrutable handwritten document images paper page: https://
arxiv.org/abs/2306.06823 State-of-the-art information extraction methods are limited by OCR errors. They work well for printed text in form-like documents, but unstructured, -

High-Fidelity Audio Compression with Improved RVQGAN
By
–
High-Fidelity Audio Compression with Improved RVQGAN paper page: https://
huggingface.co/papers/2306.06
546
… Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can -

Aladdin: Zero-Shot 3D Asset Generation from Scene Descriptions
By
–
Aladdin: Zero-Shot Hallucination of Stylized 3D Assets from Abstract Scene Descriptions paper page: https://
huggingface.co/papers/2306.06
212
… What constitutes the "vibe" of a particular scene? What should one find in "a busy, dirty city street", "an idyllic countryside", or "a crime scene in an -

Large Language Models as Tax Attorneys: Legal Capabilities Study
By
–
Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence paper page: https://
huggingface.co/papers/2306.07
075
… Better understanding of Large Language Models' (LLMs) legal analysis abilities can contribute to improving the efficiency of legal services, governing