TokenFormer, a new model architecture from @cvml_mpiinf and @PKU1898
, scales from 124M to 1.4B parameters by treating parameters as tokens, maintaining Transformer performance with lower cost. Talk to the team @haiyang73756134 @ferjadnaeem @xyongqin @janericlenssen @fedassa here!
@askalphaxiv
-

TokenFormer: Efficient Transformer Architecture Scaling from 124M to 1.4B Parameters
By
–
-

ENN Nuclear Fusion Roadmap Faces Criticism on alphaXiv
By
–
Trending on alphaXiv: criticism of a recent nuclear fusion roadmap from ENN – one of the largest clean energy distributors in China. Tune into the discussion in the links below
-
Model Swarms: Collaborative LLM Expert Algorithm via Swarm Intelligence
By
–
New from @UW and @GoogleDeepMind
: Model Swarms, a collaborative search algorithm that adapts LLM experts to single task, multi-task domains, and reward models via swarm intelligence. Talk to the team @ZifengWang315 @chl260 @YejinChoinka @tsvetshop here! https://
alphaxiv.org/abs/2410.11163
v1
… -

Agent-as-a-Judge: AI Agents Evaluating Other Agents
By
–
New from @metaai and @KAUST_News
! Introducing Agent-as-a-Judge, a new framework where AI agents evaluate other AI agents, leading to a 97% saving in time AND cost Talk to the team @MingchenZhuge
, @SchmidhuberAI
,
@tydsh
, @zechunliu
, @YoungXiong1
, Yangyang Shi, @vikasc here! -

Meta’s LongVU VideoLLM Authors Discuss Research
By
–
Authors of LongVU, the new VideoLLM from @AIatMeta
, discussing their work this week on alphaXiv! http://
alphaxiv.org/abs/2410.17434 Huge shoutout to the team @xiaoqian_shen @liuzhuang1234 @Hu_Hsu @garvinchen2 @klightlm @zechunliu @balakrishnan_vr Florian Bordes @Fanyi_Xiao @hyunwoojkim -
Make Sehun’s Dream Come True: New AI Research Breakthrough
By
–
Make Sehun's dream come true: https://
alphaxiv.org/abs/2410.14193 -
arXiv Papers Now Available in Dark Mode Reading
By
–
NEW: Read arXiv papers in dark mode! pic.twitter.com/PGTXfzWsNI
— alphaXiv (@askalphaxiv) 23 octobre 2024NEW: Read arXiv papers in dark mode!
-

OpenAI Launches MLE-Bench for ML Engineering AI Agent Evaluation
By
–
Introducing MLE-bench, a new benchmark from @OpenAI to evaluate the performance of AI agents across 75 ML engineering tasks. Excited to have the authors discussing their work here! @junshernchan @thelokasiffers @JaffeOliver @jjamesaung @danesherbs @evanon0ping @ChowdhuryNeil
-

Differential Transformers and Learnable Temperature Mechanisms
By
–
Relationship between diff transformer and using a learnable temperature 1/N
