4). Mixtures of In-Context Learners – uses subsets of demonstrations to train experts via in-context learning; given a training set, a trainable weighting function is used to combine the experts' next-token predictions…
LLMS
-

Anthropic Employees Optimistic About AI Progress Beyond Current Benchmarks
By
–
Even @AnthropicAI employees are convinced that we are not reaching any real walls. On the contrary, we should assume that today's benchmarks, which are very challenging, will be overcome in just a few years. I remain so optimistic that I believe that the concern often stems from
-
HfApiEngine improvement simplifies open-LLM agent creation
By
–
A nice PR from @BruleNaudet in transformers.agents just improved HfApiEngine. This makes the creation of open-LLM-powered agents even easier with our free Inference API!
-
GPT-5 Progress Contradicts Gary Marcus Predictions
By
–
GPT-5 is not slowing down, Gary Marcus however calls it a win for his prediction. More and more people are saying that The Information is simply wrong this time. And I agree with them completely. As I wrote in my analysis, The Information probably refers to the regular training
-
Information Error: LLMs Lack Reasoning Methods Coverage
By
–
Agreed. Looks like the information was wrong and only focuses on LLMs without reasoning methods
-

GPT-5 Development Speed: The Information vs Noam Brown
By
–
GPT-5 Development slowing down? The Information says yes, OpenAI legend Noam Brown says no! My conclusion. Here is everything we know: The Information, the well-known magazine that has long had accurate leaks and insider knowledge about OpenAI, writes in summary: "Some OpenAI
-

Jimmy Denies GPT-5 Training Diminishing Returns Article as Fake
By
–
Jimmy calls the article from the information about dimishing returns in GPT-5 / Orion training fake news. Good to know!
— Chubby♨️ (@kimmonismus) 10 novembre 2024
In about 2 hours I’ll publish a post where I’ll break down everything we know.
But until then: love and trust the apple! https://t.co/18VO4SKXrgJimmy calls the article from the information about dimishing returns in GPT-5 / Orion training fake news. Good to know! In about 2 hours I’ll publish a post where I’ll break down everything we know. But until then: love and trust the apple!
-

GPT-5 Orion Performance Falls Short of Expectations
By
–
GPT-5 not as good as expected? If our summary is correct, Orion does not seem to be doing particularly well. – the leap in comparison from GPT3 to 4 is not to be expected to the same extent with 5
– you probably have to switch to new methods
– Although synthetic data is already -
Two-Channel Audio Models with Text Pretraining Architecture
By
–
One-Channel Stack: > Trained on 20M hours of audio
> Primary checkpoint initialized from pretrained language model on 2T text tokens
> Text-pretrained model shows higher coherence in subjective evaluations Two-Channel Hertz-lm: > Predicts two quantized latents for two separate -
Hertz-VAE: 1.8B Parameter Decoder-Only Transformer Architecture
By
–
Hertz-vae: > 1.8B parameters, 8-layer decoder-only transformer
> First four layers receive latent history
> Layer 5 receives ground-truth 15-bit quantized representation during training
> Directly samples hertz-lm's next token prediction during inference
> Near-perfect at
