Open call @papla_media – if you consider Open Sourcing your model checkpoints I'd be more than happy to help from Hugging Face Hub side!
@reach_vb
-

Breakthrough: Open GPT-4o Reproduction Achieves Multimodal Capabilities
By
–
This is the closest we've been to Open GPT4o reproduction! 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 28 octobre 2024
Text + Audio + Image modalities combined, one-shot. https://t.co/13NJtqT5Bv pic.twitter.com/mg6v587f6EThis is the closest we've been to Open GPT4o reproduction! Text + Audio + Image modalities combined, one-shot.
-
Mini-Omni 2: Multimodal AI with Real-Time Voice Conversations
By
–
Mini-Omni 2 understands image, audio and text inputs all via end-to-end voice conversations with users 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 28 octobre 2024
> Understands and processes images, speech, and text
> Generates real-time speech responses
> Supports interruptions during speech
Technical Overview:
> Concatenates image,… pic.twitter.com/VrmlAWX0JjMini-Omni 2 understands image, audio and text inputs all via end-to-end voice conversations with users > Understands and processes images, speech, and text
> Generates real-time speech responses
> Supports interruptions during speech Technical Overview:
> Concatenates image, -
Free AI Spaces Available on Hugging Face Platform
By
–
It's free, you just go on http://
hf.co/spaces and click on any of them – use as you like. -
Explore Hugging Face Spaces for AI Models and Applications
By
–
Check'em out here: https://
huggingface.co/spaces -

Weekly AI Spaces: Text to Image, Video, Multilingual LLMs
By
–
Spaces of the week! – we've got Text to Image, Text to Video, Multilingual LLMs and much more!
-
Open Source AI Project Code Available for Experimentation
By
–
But, it’s open – with all the code open source, you can fiddle with it, try different prompts, things etc The final output is not at that level, but, it only gets better from here,
-
F5-TTS Audio Quality Optimization with Reference Speakers
By
–
Nice! Lmk how it goes, I’ll be back in office tomorrow will take a deeper look. Btw from my experience w/ F5-TTS – the generation quality depends quite a bit on the reference audio – might be worth checking with different speaker prompts.
-
Audio Refinement Step for AI Generation Quality Improvement
By
–
Another cool to experiment would be to add an Audio refiner as an optional Step 5 – with the sole goal to make the generation sound as good as possible. Resemble Enhance would fit in well there:
-
F5-TTS and E2 TTS: Exploring the Best Text-to-Speech Solutions
By
–
Heya! I think F5-TTS/ E2 TTS is the best atm. It’d be cool to experiment with it.