3) VeRA – In LoRA, low-rank matrices A and B are unique for each layer.
– In VeRA, A and B are frozen, random, and shared across all layers.
– Instead, it learns layer-specific scaling VECTORS (b and d) instead. Check this
VeRA: Frozen Shared Matrices with Layer-Specific Scaling Vectors
By
–
