how many params can you fit on this bad boi?
@reach_vb
-

Fish Audio TTS Model Ranks Second on TTS Arena
By
–
Fish Audio team is CRACKED AF – the model is ranking #2 overall on the TTS Arena! 🤯
— Vaibhav (VB) Srivastav (@reach_vb) 4 décembre 2024
(preliminary results, go vote nowww) https://t.co/emjNq93Fan pic.twitter.com/ffCfMoKg05Fish Audio team is CRACKED AF – the model is ranking #2 overall on the TTS Arena! (preliminary results, go vote nowww)
-
Direct model interaction now available for testing
By
–
You can play with the model directly here too:
-

FishSpeech v1.5: Open-Source Multilingual Voice Cloning Model
By
–
LETS GOOO! FishSpeech v1.5 – multilingual, zero-shot instant voice cloning, low-latency, open text to speech model 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 4 décembre 2024
> Only 500M params
> Trained on 1 MILLION hours of audio
> Supports 13 languages
> Low-latency (<150 ms)
> Open model – checkpoints on the hub 🤗
> Best part:… pic.twitter.com/kwKRBvsnabLETS GOOO! FishSpeech v1.5 – multilingual, zero-shot instant voice cloning, low-latency, open text to speech model > Only 500M params
> Trained on 1 MILLION hours of audio
> Supports 13 languages
> Low-latency ( Open model – checkpoints on the hub > Best part: -
Open Artifacts Required for Blog Post Coverage
By
–
i'll take blogposts as long as they come with open artefacts on the hub
-
User anticipates Whisper speech recognition model refresh
By
–
i was hoping for another whisper refresh actually
-
Smaller Multilingual Models Under 1B Parameters
By
–
go even smaller than 2B & multilingual – sub 1B is perfect size!
-

Autoregressive Latent Diffusion Model for Video Generation
By
–
> autoregressive latent diffusion model
— Vaibhav (VB) Srivastav (@reach_vb) 4 décembre 2024
> trained on large video datasets
> latent frames pass through an autoencoder to a transformer dynamics model
> uses a causal mask similar to LLMs
> inference involves frame-by-frame autoregressive sampling with past frames
>… https://t.co/bLiWmy1IDS pic.twitter.com/7V3vsiC6c1> autoregressive latent diffusion model > trained on large video datasets
> latent frames pass through an autoencoder to a transformer dynamics model
> uses a causal mask similar to LLMs
> inference involves frame-by-frame autoregressive sampling with past frames
> -
DeepMind Genie 2: Multimodal World Model Generates 3D Environments
By
–
DeepMind COOKED! Genie 2, a large-scale, multi-modal foundation world model! 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 4 décembre 2024
Capable of creating endless action-controllable, playable 3D environments – the future is going to be so, so wild! pic.twitter.com/BRi2djkarmDeepMind COOKED! Genie 2, a large-scale, multi-modal foundation world model! Capable of creating endless action-controllable, playable 3D environments – the future is going to be so, so wild!