Highlighting interesting contributions on our featured paper, the Differential Transformer from @ytz2024 @MSFTResearch @Tsinghua_Uni
@askalphaxiv
-
Molmo Authors Discussion Direct Engagement Opportunity Available
By
–
Discuss Molmo with the authors directly here: https://
alphaxiv.org/abs/2409.17146
v1
… -

Diff Transformer: Attention Mechanism for Long Context NLP
By
–
Introducing the Diff Transformer, which calculates attention scores as the difference between two separate softmax attention maps. This improves NLP tasks like long-context modeling, in-context retrieval, and mitigating hallucinations! Join the discussion with @ytz2024 here!
-

LlaMa-Omni: Real-time Speech LLM Model Rivals GPT-4o
By
–
Introducing LlaMa-Omni, a new model rivaling GPT-4o for real-time speech interaction with LLMs. This simultaneously generates text and speech directly from speech instructions, with a response latency as low as 226ms. Excited to have the author @Poeroz1204 discussing the work!
-

Molmo Dataset: Why Clock Images Matter for Vision Models
By
–
And finally, the reason the Molmo authors collected a dataset consisting of clocks 7/N
-
PTAS for Game Equilibria Could Redefine PPAD Complexity
By
–
For decades, researchers have debated the existence of a PTAS (polynomial-time approximation scheme) for game equilibria. In this new work from @shb20tsinghua
, the authors propose a PTAS, implying PPAD=FP. Could this overturn the belief that PPAD contains intractable problems? -

Allen AI and UW Release Open-Source Multimodal AI Models
By
–
New work from @allen_ai and @UW
: SOTA multimodal AI models that are truly open-source. The authors have open-sourced their full pipeline, including model weights (Molmo), datasets (PixMo), and training code. Excited to have @inkynumbers and team discussing their work here!



