Wohoo! DeepMind released AlphaFold 3 inference codebase, model weights and an on-demand server! AlphaFold can generate highly accurate biomolecular structure predictions containing proteins, DNA, RNA, ligands, ions, and also model chemical modifications for proteins and
@reach_vb
-
Hugging Face Hub hosts model weights with gated repository access
By
–
Congratulations on the release – Would love to help get these model weights on the Hugging Face hub. We can set up gated repositories so you have full ACL of who accesses the model repositories as well. (We do the same for Gemma repos as well, lmk if that’s helpful)
-
Fine-tune AI models with custom datasets and voice personas
By
–
You would be able to fine-tune it on your own dataset/ voice persona soon!
-
Two-Channel Audio Models with Text Pretraining Architecture
By
–
One-Channel Stack: > Trained on 20M hours of audio
> Primary checkpoint initialized from pretrained language model on 2T text tokens
> Text-pretrained model shows higher coherence in subjective evaluations Two-Channel Hertz-lm: > Predicts two quantized latents for two separate -
Hertz-VAE: 1.8B Parameter Decoder-Only Transformer Architecture
By
–
Hertz-vae: > 1.8B parameters, 8-layer decoder-only transformer
> First four layers receive latent history
> Layer 5 receives ground-truth 15-bit quantized representation during training
> Directly samples hertz-lm's next token prediction during inference
> Near-perfect at -
Hertz-lm: 6.6B Parameter Audio Language Model Released
By
–
Hertz-lm: > 6.6B parameters, 32-layer decoder-only transformer
> Context of 2048 input tokens (~4.5 mins)
> Predicts 15-bit compressed versions of hertz-codec tokens -
Hertz-codec: Advanced Convolutional Audio VAE Outperforms Competitors
By
–
Hertz-codec: > Convolutional audio VAE
> Encodes 16kHz mono speech to 8Hz latent representation at 1kbps
> 32-dim latent per 125ms frame
> Outperforms Soundstream and Encodec at 6kbps, on par with DAC at 8kbps
> 5M encoder, 95M decoder parameters -
Hertz-dev: 8.5B Parameter Open-Source Audio Model
By
–
Hertz-dev – 8.5 billion parameters, full-duplex, audio-only base model, APACHE 2.0 licensed 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 10 novembre 2024
> Trained on 20 million hours of audio
Train on any down-stream task, speech-to-speech, translation, classification, speech recognition, text-to-speech and more!
GG @si_pbc 🤗 pic.twitter.com/MlCt6njDGYHertz-dev – 8.5 billion parameters, full-duplex, audio-only base model, APACHE 2.0 licensed > Trained on 20 million hours of audio Train on any down-stream task, speech-to-speech, translation, classification, speech recognition, text-to-speech and more! GG @si_pbc