3/ MambaByte – adapts Mamba SSM to learn directly from raw bytes; bytes lead to longer sequences which autoregressive Transformers will scale poorly on; reports huge benefits related to faster inference and even outperforms subword Transformers.
MambaByte: SSM learns raw bytes, outperforms Transformers
By
–
