DORSal: Diffusion for Object-centric Representations of Scenes et al. paper page: https://
huggingface.co/papers/2306.08
068
… Recent progress in 3D scene understanding enables scalable learning of representations across large datasets of diverse scenes. As a consequence, generalization to unseen
MULTIMODAL AI
-

DORSal: Diffusion for Object-centric Scene Representations
By
–
-

Anticipatory Music Transformer: Controllable Generative Model for Temporal Point Processes
By
–
Anticipatory Music Transformer paper page: https://
huggingface.co/papers/2306.08
620
… introduce anticipation: a method for constructing a controllable generative model of a temporal point process (the event process) conditioned asynchronously on realizations of a second, correlated process (the -
VidEdit: Zero-Shot Spatially Aware Text-Driven Video Editing
By
–
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
— AK (@_akhaliq) 16 juin 2023
paper page: https://t.co/g60mxptgC1
Recently, diffusion-based generative models have achieved remarkable success for image generation and edition. However, their use for video editing still faces important… pic.twitter.com/aT67jfcETzVidEdit: Zero-Shot and Spatially Aware Text-Driven Editing paper page: https://
huggingface.co/papers/2306.08
707
… Recently, diffusion-based generative models have achieved remarkable success for image generation and edition. However, their use for video editing still faces important -
Language-Guided Music Recommendation for Videos Using Prompt Analogies
By
–
Language-Guided Music Recommendation for Video via Prompt Analogies
— AK (@_akhaliq) 16 juin 2023
paper page: https://t.co/qkQcdK2FB6
propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language. A key challenge of this problem setting… pic.twitter.com/rlHWtziEYwLanguage-Guided Music Recommendation for via Prompt Analogies paper page: https://
huggingface.co/papers/2306.09
327
… propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language. A key challenge of this problem setting -

Macaw-LLM: Multi-Modal Language Model with Image Audio Video Text
By
–
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration paper page: https://
huggingface.co/papers/2306.09
093
… Although instruction-tuned large language models (LLMs) have exhibited remarkable capabilities across various NLP tasks, their effectiveness on other data -

LOVM: Language-Only Vision Model Selection Framework
By
–
LOVM: Language-Only Vision Model Selection paper page: https://
huggingface.co/papers/2306.08
893
… Pre-trained multi-modal vision-language models (VLMs) are becoming increasingly popular due to their exceptional performance on downstream vision applications, particularly in the few- and zero-shot -
AssistGPT: Multi-modal AI Assistant with Planning and Learning
By
–
AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
— AK (@_akhaliq) 16 juin 2023
paper page: https://t.co/9KBXoqjiEx
Recent research on Large Language Models (LLMs) has led to remarkable advancements in general NLP AI assistants. Some studies have further explored the use… pic.twitter.com/Vl1aUyIgNfAssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn paper page: https://
huggingface.co/papers/2306.08
640
… Recent research on Large Language Models (LLMs) has led to remarkable advancements in general NLP AI assistants. Some studies have further explored the use -

AVIS: Autonomous Visual Information Seeking with Large Language Models
By
–
AVIS: Autonomous Visual Information Seeking with Large Language Models paper page: https://
huggingface.co/papers/2306.08
129
… In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically -
MammalNet: Large-scale Video Benchmark for Mammal Recognition and Behavior
By
–
MammalNet: A Large-scale Video Benchmark for Mammal Recognition and Behavior Understanding
— AK (@_akhaliq) 16 juin 2023
paper page: https://t.co/KFAUDnDlQO
Monitoring animal behavior can facilitate conservation efforts by providing key insights into wildlife health, population status, and ecosystem… pic.twitter.com/mkICOkdzDgMammalNet: A Large-scale Benchmark for Mammal Recognition and Behavior Understanding paper page: https://
huggingface.co/papers/2306.00
576
… Monitoring animal behavior can facilitate conservation efforts by providing key insights into wildlife health, population status, and ecosystem -
Google AI Lets You Preview Clothes On Different Models
By
–
Google's new #generativeAI lets you preview clothes on different models https://
tcrn.ch/42EVDhc via @techcrunch