Group: Model Lifecycle (bridge) | Previous group exit: inference fundamentals and interface contracts | This page exit: know when to change weights instead of prompts or retrieval
RLHF (Bridge)
Bridge page: this page answers only "how RLHF affects your application engineering decisions". Reward modeling and PPO/DPO mechanics are derived in Learn LLM (RLHF/DPO chapters).
What problem it solves
Reinforcement Learning from Human Feedback (RLHF): humans rank multiple model outputs, a reward model is trained to predict those preferences, and reinforcement learning then optimizes the model to maximize the reward score. It changes not what the model knows (that is pretraining and SFT) but how it behaves — the helpful/honest/harmless alignment behavior comes mainly from this stage.
Three steps:
- SFT: base capability from human demonstrations.
- Reward modeling: human rankings → train a reward model that predicts preference.
- RL: optimize outputs to maximize reward (PPO, or direct-alignment variants such as Direct Preference Optimization).
Why app engineers understand it but never implement it
RLHF requires large-scale human feedback data, a dedicated research team, and substantial training compute. This is the work model vendors (OpenAI, Anthropic, Google, Meta, ...) do when producing the models you consume via API. "Run RLHF ourselves" is not an option in your engineering decisions — when the need is behavioral, check the mainline toolbox first:
- Fixed tone/format → prompt engineering + few-shot examples.
- Domain knowledge → RAG.
- Refusing certain requests → input validation + output filtering (application security, see Security).
Model behavior that understanding RLHF explains
That is the real value of this page — three common "why does the model do that" cases:
- Over-refusal: alignment training amplifies boundaries; the model rejects some benign requests. Product-side response: let the user clarify and retry instead of treating refusal as terminal.
- Verbosity and hedging: "As an AI language model..." preambles. Product-side response: constrain output shape in the prompt, combined with structured output.
- Model "personality": style differences across vendors come mainly from preference-data choices at the alignment stage. Treat "does the personality match the product voice" as a selection criterion — see Evaluation (bridge).
Decision impact table
| Dimension | Notes |
|---|---|
| Data requirement | Large-scale human preference rankings (vendor-level investment; app teams do not have this) |
| Cost structure | Research-grade: feedback collection + multi-round training + evaluation (no specific figures cited here) |
| Time to effect | Model-version granularity — you only "use" a new alignment by switching model/version |
| Maintenance burden | The vendor's burden; yours is behavior regression testing after model-version upgrades |
Division of labor: SFT vs RLHF
Both bridges in one table, to defuse the "it's all training" confusion:
| Problem | Owner | What changes | Your involvement |
|---|---|---|---|
| The model doesn't know my domain/format | SFT | Knowledge and skill (weights) | A real project once data and budget exist |
| The model's behavior is wrong (tone, refusal, verbosity) | RLHF / alignment | Behavioral preference (weights) | Never a project — adapt via model selection and prompts |
| The model lacks live/private facts | RAG | Context (weights unchanged) | Everyday mainline work |
Nearly all flagship chat models go through RLHF or comparable alignment (RLAIF / Constitutional AI and the like); the differences show up mainly as each vendor's "personality" and safety boundaries — part of selection evaluation.
Deep derivations and implementation
- Reward models, PPO, DPO derivations and comparisons → the RLHF/DPO chapters of Learn LLM.
- Further reading (original papers): InstructGPT, Constitutional AI, Deep RL from Human Preferences.