github: https://
github.com/DrewThomasson/
ebook2audiobookXTTS
…
@_akhaliq
-
ebook2audiobook XTTS GitHub Repository Project
By
–
-

ebook2audiobook: Convert eBooks to Audiobooks with Chapters
By
–
ebook2audiobook Convert eBooks to audiobooks with chapters and metadata using Calibre and Coqui XTTS.
-

Hugging Face Spaces to GitHub Repository Cloning Tool
By
–
Hugging Face Spaces to GitHub Repo Enter the details to clone a Hugging Face Spaces repository and push it to a new GitHub repository. https://
huggingface.co/spaces/akhaliq
/spaces-to-github
… -

LLaVA-3D: Empowering Large Multimodal Models with 3D Awareness
By
–
LLaVA-3D
— AK (@_akhaliq) 27 septembre 2024
A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos.… pic.twitter.com/pbyiLnBhr6LLaVA-3D A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos.
-

Lotus: Diffusion-Based Visual Foundation Model for Dense Prediction
By
–
Lotus Diffusion-based Visual Foundation Model for High-quality Dense Prediction Leveraging the visual priors of pre-trained text-to-image diffusion models offers a promising solution to enhance zero-shot generalization in dense prediction tasks. However, existing methods often
-

Robot Learning Object Manipulation Through Monocular 4D Reconstruction
By
–
Robot See Robot Do
— AK (@_akhaliq) 27 septembre 2024
Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying… pic.twitter.com/TsvoduD2b0Robot See Robot Do Imitating Articulated Object Manipulation with Monocular 4D Reconstruction Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying
-

EMOVA: Language Models with Multimodal Emotions and Expression
By
–
EMOVA
— AK (@_akhaliq) 27 septembre 2024
Empowering Language Models to See, Hear and Speak with Vivid Emotions
discuss: https://t.co/QifoJVP166
GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering… pic.twitter.com/UBIT587NmlEMOVA Empowering Language Models to See, Hear and Speak with Vivid Emotions discuss: https://
huggingface.co/papers/2409.18
042
… GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering -

MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models
By
–
MaskLLM Learnable Semi-Structured Sparsity for Large Language Models discuss: https://
huggingface.co/papers/2409.17
481
… Large Language Models (LLMs) are distinguished by their massive parameter counts, which typically result in significant redundancy. This work introduces MaskLLM, a learnable -

Accelerating Long-Context LLMs with Massive Input Token Reduction
By
–
Discovering the Gems in Early Layers Accelerating Long-Context LLMs with 1000x Input Token Reduction discuss: https://
huggingface.co/papers/2409.17
422
… Large Language Models (LLMs) have demonstrated remarkable capabilities in handling long context inputs, but this comes at the cost of increased