Large language models have demonstrated a surprising range of skills and behaviors. How can we trace their source? In our new paper, we use influence functions to find training examples that contribute to a given model output.
Tracing LLM Skills Through Influence Functions on Training Data
By
–
