huge update, vote here: https://
huggingface.co/spaces/TIGER-L
ab/GenAI-Arena
…
@_akhaliq
-

Major Update Available for GenAI-Arena Voting Platform
By
–
-
Adobe Magic Fixup Enables Automatic Image Editing with Cut-and-Paste
By
–
Adobe released Magic Fixup!
— AK (@_akhaliq) 21 août 2024
local @Gradio demo: https://t.co/pCk07tfehy
enable users to edit images with simple a cut-and-paste like approach, and fixup those edits automatically. pic.twitter.com/xBC1mMAkIWAdobe released Magic Fixup! local @Gradio demo: https://
github.com/adobe-research
/MagicFixup?tab=readme-ov-file#gradio-demo
… enable users to edit images with simple a cut-and-paste like approach, and fixup those edits automatically. -
RP1M: Large-Scale Motion Dataset for Bi-Manual Dexterous Robot Piano Playing
By
–
RP1M
— AK (@_akhaliq) 21 août 2024
A Large-Scale Motion Dataset for Piano Playing with Bi-Manual Dexterous Robot Hands
discuss: https://t.co/7eMw6eOOFR
It has been a long-standing research goal to endow robot hands with human-level dexterity. Bi-manual robot piano playing constitutes a task that combines… pic.twitter.com/T0KPLeguL8RP1M A Large-Scale Motion Dataset for Piano Playing with Bi-Manual Dexterous Robot Hands discuss: https://
huggingface.co/papers/2408.11
048
… It has been a long-standing research goal to endow robot hands with human-level dexterity. Bi-manual robot piano playing constitutes a task that combines -

ShapeSplat: Large-scale Gaussian Splats Dataset and Self-Supervised Pretraining
By
–
ShapeSplat A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining' discuss: https://
huggingface.co/papers/2408.10
906
… 3D Gaussian Splatting (3DGS) has become the de facto method of 3D representation in many vision tasks. This calls for the 3D understanding directly in -

MagicDec: Speculative Decoding Breaks Latency-Throughput Tradeoff
By
–
MagicDec Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding discuss: https://
huggingface.co/papers/2408.11
049
… Large Language Models (LLMs) have become more prevalent in long-context applications such as interactive chatbots, document analysis, and -

MegaFusion: Higher-Resolution Image Generation Without Tuning
By
–
MegaFusion Extend Diffusion Models towards Higher-resolution Image Generation without Further Tuning discuss: https://
huggingface.co/papers/2408.11
001
… Diffusion models have emerged as frontrunners in text-to-image generation for their impressive capabilities. Nonetheless, their fixed image -

Transfusion: Multi-Modal Model for Token Prediction and Image Diffusion
By
–
Transfusion Predict the Next Token and Diffuse Images with One Multi-Modal Model discuss: https://
huggingface.co/papers/2408.11
039
… We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function -
Audio Match Cutting: Finding Matching Audio Transitions in Videos
By
–
Audio Match Cutting
— AK (@_akhaliq) 21 août 2024
Finding and Creating Matching Audio Transitions in Movies and Videos
discuss: https://t.co/bMy5MSiGeY
A "match cut" is a common video editing technique where a pair of shots that have a similar composition transition fluidly from one to another. Although… pic.twitter.com/9VRs0Fx9nQAudio Match Cutting Finding and Creating Matching Audio Transitions in Movies and Videos discuss: https://
huggingface.co/papers/2408.10
998
… A "match cut" is a common video editing technique where a pair of shots that have a similar composition transition fluidly from one to another. Although -

MambaEVT: Event Stream Visual Object Tracking with State Space Model
By
–
MambaEVT Event Stream based Visual Object Tracking using State Space Model discuss: https://
huggingface.co/papers/2408.10
487
… Event camera-based visual tracking has drawn more and more attention in recent years due to the unique imaging principle and advantages of low energy consumption, high -

New Guide: Enhance CogVideoX Generated Videos with VEnhancer
By
–
New, Enhance CogVideoX Generated Videos with VEnhancer guide https://
github.com/THUDM/CogVideo
/tree/main/tools/venhancer
…