AI Dynamics

Global AI News Aggregator

About

TokenFormer: Efficient Transformer Architecture Scaling from 124M to 1.4B Parameters

TokenFormer, a new model architecture from @cvml_mpiinf and @PKU1898
, scales from 124M to 1.4B parameters by treating parameters as tokens, maintaining Transformer performance with lower cost. Talk to the team @haiyang73756134 @ferjadnaeem @xyongqin @janericlenssen @fedassa here!

→ View original post on X — @askalphaxiv