On the Exploitability of Instruction Tuning paper page: https://
huggingface.co/papers/2306.17
194
… Instruction tuning is an effective technique to align large language models (LLMs) with human intents. In this work, we investigate how an adversary can exploit instruction tuning by injecting
Exploitability of Instruction Tuning in Large Language Models
By
–
