Special thanks to: – @failspy for his notebook and abliterated models
– Arditi et al. (cc @NeelNanda5
) for the excellent "Refusal in LLMs" blog post (
https://
lesswrong.com/posts/jGuXSZgv
6qfdhMCuJ/refusal-in-llms-is-mediated-by-a-single-direction
…)
– All the LLM mergers and fine-tuners for the source models
– Charles Goddard and @arcee_ai for MergeKit
LLM Interpretability and Refusal Mechanisms Research Acknowledgments
By
–