7/ DoReMi – trains a small proxy model over domains to produce domain weights without knowledge of downstream tasks; it resamples a dataset with the domain weights which allows using a 280M proxy model to train an 8B model (30x larger) more efficiently.
LLMS
-
TinyStories: Efficient Language Models for Story Generation
By
–
6/ TinyStories – uses a synthetic dataset of short stories to train and evaluate LMs that are much smaller than SoTA models but can produce fluent and consistent stories with several paragraphs, and demonstrate reasoning capabilities.
-
Top ML Papers of the Week: DragGAN, CodeT5+, Med-PaLM 2
By
–
Top ML Papers of the Week (May 15 – 21): – DragGAN
– CodeT5+
– StructGPT
– Med-PaLM 2
– Symbol Tuning
– Evidence of Meaning in LLMs
… -
Constitutional AI: Explicit Values Framework for Language Models
By
–
Constitutional AI is the act of giving a large language model “explicit values determined by a constitution, rather than values determined implicitly via large-scale human feedback.”
-
Data Cleanup Work Required for GPT Equivalence Often Underestimated
By
–
Also your comment trivialises the sheer amount of data and cleanup work someone has to do to get to GPT equivalence.
-
Fine-tuning Creates GPT-Equivalent LLMs, Not Just Wrappers
By
–
It’s not a wrapper. It’s fine tuning. And the end result is an LLM equivalent to GPT. ChatGPT stands on the shoulders of many giants it did not build either.
-
Fine-tuning Llama to match GPT-3.5 performance
By
–
Fine tune llama to near perfectly match GPT -3.5
-
LoRA and Small Base Models as Cost-Effective AI Solutions
By
–
Lora and small base models to the rescue In all seriousness it isn’t THAT expensive to get a bunch of A100s..
-
Can an Indian team build GPT-3.5 equivalent LLM
By
–
Do you think an Indian team sitting in India can build an LLM equivalent to GPT-3.5?
