Hacker News for AI research PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions Hot right now
@_akhaliq
-
Hugging Face Daily Papers: Stay Updated with AI Research
By
–
New Blog! Stay ahead in AI with Hugging Face’s Daily Papers! Join 12k+ subscribers exploring 3,700+ top research papers daily. Features include: Claim your work Submit & discover papers Chat with authors Access models & demos Upvote favorites
-

Google Zero-shot Cross-lingual Voice Transfer for TTS
By
–
Google presents Zero-shot Cross-lingual Voice Transfer for TTS discuss: https://
huggingface.co/papers/2409.13
910
… In this paper, we introduce a zero-shot Voice Transfer (VT) module that can be seamlessly integrated into a multi-lingual Text-to-speech (TTS) system to transfer an individual's -

MaterialFusion: Enhancing Inverse Rendering with Material Diffusion Priors
By
–
MaterialFusion
— AK (@_akhaliq) 24 septembre 2024
Enhancing Inverse Rendering with Material Diffusion Priors
discuss: https://t.co/Es6DJE83P3
Recent works in inverse rendering have shown promise in using multi-view images of an object to recover shape, albedo, and materials. However, the recovered components… pic.twitter.com/RldZpMpcp9MaterialFusion Enhancing Inverse Rendering with Material Diffusion Priors discuss: https://
huggingface.co/papers/2409.15
273
… Recent works in inverse rendering have shown promise in using multi-view images of an object to recover shape, albedo, and materials. However, the recovered components -

Nvidia MaskedMimic: Unified Physics-Based Character Control
By
–
Nvidia presents MaskedMimic
— AK (@_akhaliq) 24 septembre 2024
Unified Physics-Based Character Control Through Masked Motion Inpainting
discuss: https://t.co/Q2Tx01NP2I
Crafting a single, versatile physics-based controller that can breathe life into interactive characters across a wide spectrum of scenarios… pic.twitter.com/E3lADK3FcZNvidia presents MaskedMimic Unified Physics-Based Character Control Through Masked Motion Inpainting discuss: https://
huggingface.co/papers/2409.14
393
… Crafting a single, versatile physics-based controller that can breathe life into interactive characters across a wide spectrum of scenarios -

Self-Supervised Audio-Visual Soundscape Stylization Techniques
By
–
Self-Supervised Audio-Visual Soundscape Stylization
— AK (@_akhaliq) 24 septembre 2024
discuss: https://t.co/mEqithUXlD
Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input… pic.twitter.com/151CoB2BaMSelf-Supervised Audio-Visual Soundscape Stylization discuss: https://
huggingface.co/papers/2409.14
340
… Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input -
Try out Luma AI Dream Machine API app
By
–
app to try out API here: https://
huggingface.co/spaces/lumaai/
dream-machine
… -

PaliGemma Android Implementation with Hugging Face Gradio API
By
–
PaliGemma Android HF
— AK (@_akhaliq) 23 septembre 2024
github: https://t.co/A0oBkmj39S
This repository is an implementation of inferring the PaliGemma Vision Language Model on Android using Hugging Face-Gradio Client API for tasks such as zero-shot object detection, image captioning and visual… pic.twitter.com/rfB2kR8ijhPaliGemma Android HF github: https://
github.com/NSTiwari/PaliG
emma-Android-HF
… This repository is an implementation of inferring the PaliGemma Vision Language Model on Android using Hugging Face-Gradio Client API for tasks such as zero-shot object detection, image captioning and visual -

Tuning-Free Personalized Image Generation from Meta AI
By
–
Hacker News for AI research papers Imagine yourself: Tuning-Free Personalized Image Generation from @AIatMeta is hot right now
-
PDF2Audio Demo – Convert Documents to Speech
By
–
demo: https://
huggingface.co/spaces/lamm-mi
t/PDF2Audio
…
