Skip to content

Group: Model Lifecycle (bridge) | Previous group exit: inference fundamentals and interface contracts | This page exit: know when to change weights instead of prompts or retrieval

RLHF (Bridge) ​

Bridge page: this page answers only "how RLHF affects your application engineering decisions". Reward modeling and PPO/DPO mechanics are derived in Learn LLM (RLHF/DPO chapters).

What problem it solves ​

Reinforcement Learning from Human Feedback (RLHF): humans rank multiple model outputs, a reward model is trained to predict those preferences, and reinforcement learning then optimizes the model to maximize the reward score. It changes not what the model knows (that is pretraining and SFT) but how it behaves — the helpful/honest/harmless alignment behavior comes mainly from this stage.

Three steps:

  1. SFT: base capability from human demonstrations.
  2. Reward modeling: human rankings → train a reward model that predicts preference.
  3. RL: optimize outputs to maximize reward (PPO, or direct-alignment variants such as Direct Preference Optimization).

Why app engineers understand it but never implement it ​

RLHF requires large-scale human feedback data, a dedicated research team, and substantial training compute. This is the work model vendors (OpenAI, Anthropic, Google, Meta, ...) do when producing the models you consume via API. "Run RLHF ourselves" is not an option in your engineering decisions — when the need is behavioral, check the mainline toolbox first:

  • Fixed tone/format → prompt engineering + few-shot examples.
  • Domain knowledge → RAG.
  • Refusing certain requests → input validation + output filtering (application security, see Security).

Model behavior that understanding RLHF explains ​

That is the real value of this page — three common "why does the model do that" cases:

  1. Over-refusal: alignment training amplifies boundaries; the model rejects some benign requests. Product-side response: let the user clarify and retry instead of treating refusal as terminal.
  2. Verbosity and hedging: "As an AI language model..." preambles. Product-side response: constrain output shape in the prompt, combined with structured output.
  3. Model "personality": style differences across vendors come mainly from preference-data choices at the alignment stage. Treat "does the personality match the product voice" as a selection criterion — see Evaluation (bridge).

Decision impact table ​

DimensionNotes
Data requirementLarge-scale human preference rankings (vendor-level investment; app teams do not have this)
Cost structureResearch-grade: feedback collection + multi-round training + evaluation (no specific figures cited here)
Time to effectModel-version granularity — you only "use" a new alignment by switching model/version
Maintenance burdenThe vendor's burden; yours is behavior regression testing after model-version upgrades

Division of labor: SFT vs RLHF ​

Both bridges in one table, to defuse the "it's all training" confusion:

ProblemOwnerWhat changesYour involvement
The model doesn't know my domain/formatSFTKnowledge and skill (weights)A real project once data and budget exist
The model's behavior is wrong (tone, refusal, verbosity)RLHF / alignmentBehavioral preference (weights)Never a project — adapt via model selection and prompts
The model lacks live/private factsRAGContext (weights unchanged)Everyday mainline work

Nearly all flagship chat models go through RLHF or comparable alignment (RLAIF / Constitutional AI and the like); the differences show up mainly as each vendor's "personality" and safety boundaries — part of selection evaluation.

Deep derivations and implementation ​

Built for frontend engineers · Powered by VitePress