Align language models with human preferences using Direct Preference Optimization — no reward model, no RL loop.