Background Prompting for Improved Object Depth paper page: https://
huggingface.co/papers/2306.05
428
… Estimating the depth of objects from a single image is a valuable task for many vision, robotics, and graphics applications. However, current methods often fail to produce accurate depth for
@_akhaliq
-

Background Prompting for Improved Object Depth Estimation
By
–
-
Test-Time Optimization for Dense Motion Tracking in Videos
By
–
Tracking Everything Everywhere All at Once
— AK (@_akhaliq) 9 juin 2023
paper page: https://t.co/UwLxRYPGvb
present a new test-time optimization method for estimating dense and long-range motion from a video sequence. Prior optical flow or particle video tracking algorithms typically operate within limited… pic.twitter.com/3ryHUA4c9nTracking Everything Everywhere All at Once paper page: https://
huggingface.co/papers/2306.05
422
… present a new test-time optimization method for estimating dense and long-range motion from a video sequence. Prior optical flow or particle video tracking algorithms typically operate within limited -

Scaling Spherical CNNs: Spectral Domain Convolutions
By
–
Scaling Spherical CNNs paper page: https://
huggingface.co/papers/2306.05
420
… Spherical CNNs generalize CNNs to functions on the sphere, by using spherical convolutions as the main linear operation. The most accurate and efficient way to compute spherical convolutions is in the spectral domain -

R-MAE: Regions Meet Masked Autoencoders for Vision Tasks
By
–
R-MAE: Regions Meet Masked Autoencoders paper page: https://
huggingface.co/papers/2306.05
411
… Vision-specific concepts such as "region" have played a key role in extending general machine learning frameworks to tasks like object detection. Given the success of region-based detectors for -

LU-NeRF: Scene and Pose Estimation Using Local Unposed NeRFs
By
–
LU-NeRF: Scene and Pose Estimation by Synchronizing Local Unposed NeRFs paper page: https://
huggingface.co/papers/2306.05
410
… A critical obstacle preventing NeRF models from being deployed broadly in the wild is their reliance on accurate camera poses. Consequently, there is growing interest in -
Grounded Text-to-Image Synthesis with Attention Refocusing
By
–
Grounded Text-to-Image Synthesis with Attention Refocusing
— AK (@_akhaliq) 9 juin 2023
paper page: https://t.co/3DfgBmfB2I
Driven by scalable diffusion models trained on large-scale paired text-image datasets, text-to-image synthesis methods have shown compelling results. However, these models still fail… pic.twitter.com/nQRFzGoSLjGrounded Text-to-Image Synthesis with Attention Refocusing paper page: https://
huggingface.co/papers/2306.05
427
… Driven by scalable diffusion models trained on large-scale paired text-image datasets, text-to-image synthesis methods have shown compelling results. However, these models still fail -

MIMIC-IT: Multi-Modal In-Context Instruction Tuning for Vision-Language Tasks
By
–
MIMIC-IT: Multi-Modal In-Context Instruction Tuning paper page: https://
huggingface.co/papers/2306.05
425
… High-quality instructions and responses are essential for the zero-shot performance of large language models on interactive natural language tasks. For interactive vision-language tasks -

Modular Visual Question Answering Framework Using Code Generation
By
–
Modular Visual Question Answering via Code Generation paper page: https://
huggingface.co/papers/2306.05
392
… present a framework that formulates visual question answering as modular code generation. In contrast to prior work on modular approaches to VQA, our approach requires no additional -

MusicGen: Simple and Controllable Music Generation with Language Models
By
–
Simple and Controllable Music Generation paper page: https://
huggingface.co/papers/2306.05
284
… introduce MusicGen, a single Language Model (LM) that operates over several streams of compressed discrete music representation, i.e., tokens. Unlike prior work, MusicGen is comprised of a single-stage -

Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Experts
By
–
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts paper page: https://
huggingface.co/papers/2306.04
845
… Weight-sharing supernet has become a vital component for performance estimation in the state-of-the-art (SOTA) neural architecture