If you want to try blending two styles in one environment with SD3.5 Large, here’s the base prompt we used, just swap out the characters and background to make it your own: A high-quality selfie of a real man in a casual Halloween sweater, standing at a Halloween party with a
CREATIVE AI
-
MaskGCT Open Source Text-to-Speech Model Achieves New SoTA
By
–
Fuck yeah! MaskGCT – New open SoTA Text to Speech model! 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 30 octobre 2024
> Zero-shot voice cloning
> Emotional TTS
> Trained on 100K hours of data
> Long form synthesis
> Variable speed synthesis
> Bilingual – Chinese & English
> Available on Hugging Face
Fully non-autoregressive… pic.twitter.com/CAUX6cTiAGFuck yeah! MaskGCT – New open SoTA Text to Speech model! > Zero-shot voice cloning
> Emotional TTS
> Trained on 100K hours of data
> Long form synthesis
> Variable speed synthesis
> Bilingual – Chinese & English
> Available on Hugging Face Fully non-autoregressive -
Medium-sized image generation model surpasses competitors in quality
By
–
This model delivers best-in-class image generation for its size, with advanced multi-resolution capabilities. It surpasses other medium-sized models with its prompt adherence and image quality, making it a top choice for efficient, high-quality performance. You can download the
-

Stable Diffusion 3.5 Medium: Free Open Model for Consumer Hardware
By
–
Stable Diffusion 3.5 Medium is here – this open model is free for both commercial and non-commercial use. With 2.5 billion parameters, this model is designed to run “out of the box” on consumer hardware, even on a toaster! (1/3)
-

Meta Introduces MarDini for Large-Scale Video Generation
By
–
Meta presents MarDini
— AK (@_akhaliq) 29 octobre 2024
Masked Autoregressive Diffusion for Video Generation at Scale pic.twitter.com/h2u0OnroF5Meta presents MarDini Masked Autoregressive Diffusion for Generation at Scale
-

FasterCache: Training-Free Video Diffusion Model Acceleration
By
–
FasterCache
— AK (@_akhaliq) 28 octobre 2024
Training-Free Video Diffusion Model Acceleration with High Quality pic.twitter.com/qHu9Raj1pKFasterCache Training-Free Diffusion Model Acceleration with High Quality
-
F5-TTS Audio Quality Optimization with Reference Speakers
By
–
Nice! Lmk how it goes, I’ll be back in office tomorrow will take a deeper look. Btw from my experience w/ F5-TTS – the generation quality depends quite a bit on the reference audio – might be worth checking with different speaker prompts.
-
Audio Refinement Step for AI Generation Quality Improvement
By
–
Another cool to experiment would be to add an Audio refiner as an optional Step 5 – with the sole goal to make the generation sound as good as possible. Resemble Enhance would fit in well there:
-
F5-TTS and E2 TTS: Exploring the Best Text-to-Speech Solutions
By
–
Heya! I think F5-TTS/ E2 TTS is the best atm. It’d be cool to experiment with it.