This work is just the beginning. We hope to analyze the interactions between pretraining and finetuning, and combine influence functions with mechanistic interpretability to reverse engineer the associated circuits. You can read more on our blog:
Analyzing Pretraining Finetuning Interactions and Neural Circuits
By
–