Alignment: RLHF and beyond
Pre trained models predict the next token. They are not aligned with human preferences — they will complete harmful prompts, hallucinate confidently, and reproduce biases in the training data. Alignment is the process of steering model behavior toward human intent.