(10/12) Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model
Authors: Yeskendir Koishekenov, Vassilina Nikoulina, Alexandre Berard
LLMS
-

Memory-Efficient NLLB-200: Language-Specific Expert Pruning
By
–
-

SparseGPT: One-Shot Pruning of Massive Language Models
By
–
(8/12) SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Authors: @elias_frantar
, @DAlistarh -

Large Language Models as Reasoning Teachers
By
–
(7/12) Large Language Models Are Reasoning Teachers
Authors: @itsnamgyu
, Laura Schmid, Se-Young Yun -

Teaching Small Language Models to Reason
By
–
(6/12) Teaching Small Language Models to Reason
Authors: Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, @ericmalmi
, Aliaksei Severyn -

A Watermark for Large Language Models
By
–
(4/12) A Watermark for Large Language Models
Authors: @jwkirchenbauer
, @jonasgeiping
, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein -

Hungry Hungry Hippos: Language Modeling with State Space Models
By
–
(3/12) Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Authors: @tri_dao
, @realDanFu
, Khaled K. Saab, @ai_with_brains, Atri Rudra, Christopher Ré @HazyResearch @SnorkelAI -

OPT-IML: Scaling Language Model Instruction Meta Learning
By
–
(2/12) OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
Authors: Srinivasan Iyer, @VictoriaLinML
, Ramakanth Pasunuru, @tbmihaylov
, @simigd
, Ping Yu, et al. -
Top NLP Papers January 2023: Language Models and Text Generation
By
–
(1/12) Stay ahead of the curve with the top NLP papers of January 2023. From language models to text generation, this curated list by @forai_ml has got you covered on the latest advancements in language AI
This social post has been generated with Cohere https://
txt.cohere.ai/top-natural-la
nguage-processing-nlp-papers-of-january-2023
… -
Companies Spending Billions on RLHF Training for LLMs
By
–
we're starting to see top companies spend the same amount on RLHF and compute in training ChatGPT-like LLMs for example, OpenAI hired >1000 devs to RLHF their code models crazy—but soon companies will start spending $ hundreds of Ms or $ billions on RLHF, just as w/compute