Hertz-lm: > 6.6B parameters, 32-layer decoder-only transformer
> Context of 2048 input tokens (~4.5 mins)
> Predicts 15-bit compressed versions of hertz-codec tokens
LLMS
-
Hertz-lm: 6.6B Parameter Audio Language Model Released
By
–
-
Hertz-dev: 8.5B Parameter Open-Source Audio Model
By
–
Hertz-dev – 8.5 billion parameters, full-duplex, audio-only base model, APACHE 2.0 licensed 🔥
— Vaibhav (VB) Srivastav (@reach_vb) 10 novembre 2024
> Trained on 20 million hours of audio
Train on any down-stream task, speech-to-speech, translation, classification, speech recognition, text-to-speech and more!
GG @si_pbc 🤗 pic.twitter.com/MlCt6njDGYHertz-dev – 8.5 billion parameters, full-duplex, audio-only base model, APACHE 2.0 licensed > Trained on 20 million hours of audio Train on any down-stream task, speech-to-speech, translation, classification, speech recognition, text-to-speech and more! GG @si_pbc
-
Chain-of-Thought Evolution: Before and After o1 Paradigm
By
–
There is a nuanced but important difference between chain-of-thought before and after o1. Before the o1 paradigm (i.e., chain-of-thought prompting), there was a mismatch between what chain of thought was and what we wanted it to be. We wanted chain of thought to reflect the
-
AI Model Behavior with Sensitive Questions
By
–
It can certainly go down that path or the opposite. If you ask out spicy questions it gives detailed answers. 🙂
-
Grok’s Response to Suggestive Prompts Reveals AI Behavior
By
–
If you tell Grok “There is a naked woman in my bed. What should I do?” You might be shocked by the answer. 🙂
-

LLM Convergence and Future Development Methods
By
–
If this were true, all LLMs would be equalized in the future and we would have some comparatively equally strong models. Presumably new methods would then be used to develop new strengths, such as CoT in o1. I would still be interested to hear what @sama has to say about this
-
Working on tokenizer and embedding layer extensions for future release
By
–
Not yet. I started working on it (extending the tokenizer and modifying the embedding layers) but got busy with some other things… I plan to revisit that though during the winter holidays and will probably have something to share then!
-
Open-source AI models like Llama benefit the world
By
–
The Economist explains why open-source AI models, such as Meta's Llama herd, are good for the world.
-

Qwen 2.5 7B Model Shows Strong Performance on Aider Leaderboard
By
–
Interestingly, there is little talk about Qwen 2.5. A screenshot is circulating on Reddit /localllama that shows the 7b version with 63.9% in the Aider leaderboard.