I think so. In their 3.35B Tiny Aya report, they say "We use parallel Transformer blocks, which lead to a signifi-
cant improvement in training efficiency without hurting model quality."
RESEARCH
-
Technical training efficiency insights from the Tiny Aya model report
By
–
-

Technical Analysis of Parallel Block Design in LLM Architectures
By
–
It's been *almost* a bit quiet around LLM architecture releases in the past two weeks Interesting tidbit is the parallel block design. Via the Cmd-A the tech report "equivalent performance but significant improvement in throughput compared to the vanilla transformer block."
-

Anthropic gets up to GB200 capacity at Colossus 2 with SpaceX
By
–

Anthropic SpaceX Anthropic is getting up to GB200 of capacity in Colossus 2 in June as a part of the expanded agreement with SpaceX. Partnerships are huge unlocks
-

Dynamic Fine-Tuning (DFT) combines SFT loss with forward KL penalty
By
–
This is so neat! Dynamic Fine-Tuning (DFT) reweights the SFT loss by the model's own token probability, which creates a feedback loop. So they added forward KL to penalize any token the base finds likely, but the policy has pushed toward zero probability. DFT and forward KL
-
Inquiry into AI model training methodology and data augmentation
By
–
i would be pretty interested to know if anything other than scale had changed and want to know about the training, whether there was a lot of symbolically generated augmented data, etc
-
AI Models Enhance Robot Control, Accelerating Robotics Adoption
By
–
It is truly mind-blowing how capable Claude and Codex now are at configuring and controlling robots. My own experiments (with a lerobot 101) point to a fascinating new trend in robotics that could accelerate adoption. Many thanks to @Ken_Goldberg and @spencerhuang_ for their
-
Inquiry into the architectural basis of recent AI math breakthroughs
By
–
is the new math result neurosymbolic with Lean, harnesses etc or a pure LLM?
-
Neurosymbolic systems vs LLMs for mathematical reasoning
By
–
i am betting it was a neurosymbolic system rather than a pure LLM, and i have already said that’s a route to doing well in math. have not seen the details
-
Debating the Role of Neurosymbolic Systems in Mathematical AI
By
–
how much you want to bet that symbolic tools such as lean were involved and that this was not a pure LLM? i have said numerous times that neurosymbolic systems do well on math: pretty sure this was one.
-
General-purpose AI model solves major open problem in mathematics
By
–
a general-purpose model solved a major open problem in mathematics. we'll be saying this a lot over the coming years, but this is a kinda big milestone. i'm very excited for AI to greatly extend our understanding of the world, but still, i have complicated feelings today.