Anticipatory Music Transformer paper page: https://
huggingface.co/papers/2306.08
620
… introduce anticipation: a method for constructing a controllable generative model of a temporal point process (the event process) conditioned asynchronously on realizations of a second, correlated process (the
GENERATIVE AI
-

Anticipatory Music Transformer: Controllable Generative Model for Temporal Point Processes
By
–
-
VidEdit: Zero-Shot Spatially Aware Text-Driven Video Editing
By
–
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
— AK (@_akhaliq) 16 juin 2023
paper page: https://t.co/g60mxptgC1
Recently, diffusion-based generative models have achieved remarkable success for image generation and edition. However, their use for video editing still faces important… pic.twitter.com/aT67jfcETzVidEdit: Zero-Shot and Spatially Aware Text-Driven Editing paper page: https://
huggingface.co/papers/2306.08
707
… Recently, diffusion-based generative models have achieved remarkable success for image generation and edition. However, their use for video editing still faces important -

Large-scale Language Models for ASR Rescoring on Long-form Video Data
By
–
Large-scale Language Model Rescoring on Long-form Data paper page: https://
huggingface.co/papers/2306.08
133
… In this work, we study the impact of Large-scale Language Models (LLM) on Automated Speech Recognition (ASR) of YouTube videos, which we use as a source for long-form ASR. We demonstrate -

QR Code AI Art Generator with ControlNet Models
By
–
QR Code AI Art Generator demo: https://
huggingface.co/spaces/hugging
face-projects/QR-code-AI-art-generator
… These ControlNet models have been trained on a large dataset of 150,000 QR code + QR code artwork couples. They provide a solid foundation for generating QR code-based artwork that is aesthetically pleasing, while still -
Language-Guided Music Recommendation for Videos Using Prompt Analogies
By
–
Language-Guided Music Recommendation for Video via Prompt Analogies
— AK (@_akhaliq) 16 juin 2023
paper page: https://t.co/qkQcdK2FB6
propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language. A key challenge of this problem setting… pic.twitter.com/rlHWtziEYwLanguage-Guided Music Recommendation for via Prompt Analogies paper page: https://
huggingface.co/papers/2306.09
327
… propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language. A key challenge of this problem setting -

KoLA: Benchmarking World Knowledge of Large Language Models
By
–
KoLA: Carefully Benchmarking World Knowledge of Large Language Models paper page: https://
huggingface.co/papers/2306.09
296
… The unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we -

ChessGPT: Bridging Policy Learning and Language Modeling
By
–
ChessGPT: Bridging Policy Learning and Language Modeling paper page: https://
huggingface.co/papers/2306.09
200
… When solving decision-making tasks, humans typically depend on information from two key sources: (1) Historical policy data, which provides interaction replay from the environment, and -

Macaw-LLM: Multi-Modal Language Model with Image Audio Video Text
By
–
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration paper page: https://
huggingface.co/papers/2306.09
093
… Although instruction-tuned large language models (LLMs) have exhibited remarkable capabilities across various NLP tasks, their effectiveness on other data -

Create Your Own AI-Generated QR Codes
By
–
make your own AI generated QR codes https://
huggingface.co/spaces/hugging
face-projects/AI-QR-code-generator
… -

LOVM: Language-Only Vision Model Selection Framework
By
–
LOVM: Language-Only Vision Model Selection paper page: https://
huggingface.co/papers/2306.08
893
… Pre-trained multi-modal vision-language models (VLMs) are becoming increasingly popular due to their exceptional performance on downstream vision applications, particularly in the few- and zero-shot