The sense used in the InstructGPT paper is good — a model is aligned when it does what its designers want. Instruction tuning is the canonical form of LLM alignment, but earlier methods like filtering pre-train data of undesired content count too.
AI Alignment and Instruction Tuning Discussed in InstructGPT Context
By
–