5/ Fine-Tuning Language Models with Just Forward Passes – proposes a memory-efficient zeroth-order optimizer and a corresponding SGD algorithm to finetune large LMs with the same memory footprint as inference.
Memory-Efficient Fine-Tuning Large Language Models
By
–
