Anyone got any guesses as to why high-rank lora might train worse than either of low-rank low *and* full finetune? Must be something about the SGD training dynamics, right? https://
x.com/aicrumb/status
/aicrumb/status/1724911477725778148
…
@jeremyphoward
-
Why High-Rank LoRA Trains Worse Than Low-Rank and Full Finetuning
By
–
-
Training Dynamics: Multiple vs Single Linear Layers
By
–
I saw some paper mentioned on twitter last week about the different training dynamics between multiple linears layers vs single layers — even although they are mathematically the same, the training is different. I can't find it now though 🙁
-
Open Source Licensing Critical for Research Project Success
By
–
That's a totally fair perspective. Just two thoughts, if I may be so bold: – Use the term "source available", not "open source"
– Consider an open source license in the future — MSFT Research's history is littered with projects that failed due to closed licenses! -
Open License Apache 2 Unlocks Project Potential
By
–
I'd *love* to see this great work reach its full potential — a fully open license such as Apache 2 would unlock that, I think.
-
Microsoft Research License Restrictions Impact Academic AI Innovation
By
–
This looks very exciting Sebastien – congrats! Any chance we might see the license open up in the future? Currently it's research-only. MSFT has released most cool academic stuff as research-only for the last 25 years or so, and it's really limited its impact IMO.
-
Test Set Leakage Accusations Require Solid Evidence
By
–
Jay, that's really not cool to throw around such accusations unless you're *very* sure you can back them up. If you think there might be some evidence of test set leakage, publish that evidence, and let's discuss it. I've seen no compelling evidence myself so far.
-
Microsoft Confuses Open Source with Source Available
By
–
IIRC in the past MSFT has confused "open source" with "source available" in their messaging. I guess that's what's happening here too.
-
Understanding Max’s Motivations Behind AI Statements
By
–
The thing I'm confused about is that Max knows all this, but says these things anyway. I wonder what his motivation is.
-
Normalization Layers in Deep Learning Model Experiments
By
–
Thanks for sharing this interesting work. What kind of normalisation layers did you model use for these experiments?