JT-CV-9B
— AK (@_akhaliq) 25 septembre 2024
High-quality Text-to-video Models pic.twitter.com/cmVLmSMSop
JT-CV-9B High-quality Text-to-video Models

By
–
JT-CV-9B
— AK (@_akhaliq) 25 septembre 2024
High-quality Text-to-video Models pic.twitter.com/cmVLmSMSop
JT-CV-9B High-quality Text-to-video Models

By
–
MaskBit
— AK (@_akhaliq) 25 septembre 2024
Embedding-free Image Generation via Bit Tokens
1. We study the key ingredients of recent closed-source VQGAN tokenizers and develop a publicly available, reproducible, and high-performing VQGAN model, called VQGAN+, achieving a significant improvement of 6.28 rFID over… pic.twitter.com/NoTpp86WDZ
MaskBit Embedding-free Image Generation via Bit Tokens 1. We study the key ingredients of recent closed-source VQGAN tokenizers and develop a publicly available, reproducible, and high-performing VQGAN model, called VQGAN+, achieving a significant improvement of 6.28 rFID over

By
–
Alibaba presents MIMO
— AK (@_akhaliq) 25 septembre 2024
Controllable Character Video Synthesis with Spatial Decomposed Modeling
Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D… pic.twitter.com/sAozQvggNz
Alibaba presents MIMO Controllable Character Synthesis with Spatial Decomposed Modeling Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D

By
–
Making Text Embedders Few-Shot Learners discuss: https://
huggingface.co/papers/2409.15
700
… Large language models (LLMs) with decoder-only architectures demonstrate remarkable in-context learning (ICL) capabilities. This feature enables them to effectively handle both familiar and novel tasks by

By
–
Gen2Act
— AK (@_akhaliq) 25 septembre 2024
Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
discuss: https://t.co/2lTOeBcQqz
How can robot manipulation policies generalize to novel tasks involving unseen object types and new motions? In this paper, we provide a solution in… pic.twitter.com/KttbmLmzO1
Gen2Act Human Generation in Novel Scenarios enables Generalizable Robot Manipulation discuss: https://
huggingface.co/papers/2409.16
283
… How can robot manipulation policies generalize to novel tasks involving unseen object types and new motions? In this paper, we provide a solution in

By
–
try out Gemini-1.5-flash-8b-exp-0924 chatbot on @huggingface
: https://
huggingface.co/spaces/akhaliq
/gemini-1.5-flash-8b-exp-0924
…

By
–
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming demo is out by @freddy_alfonso_ demo: https://
huggingface.co/spaces/gradio/
omni-mini
…

By
–
Whisper-WebUI github: https://
github.com/jhj0517/Whispe
r-WebUI
… A Gradio-based browser interface for Whisper. You can use it as an Easy Subtitle Generator!

By
–
Phantom demo is out Phantom is super efficient 0.5B, 1.8B, 3.8B, and 7B size Large Language and Vision Models built on new propagation strategy demo: https://
huggingface.co/spaces/BK-Lee/
Phantom
…