Skip to content

Group: Model Lifecycle (bridge) | Previous group exit: inference fundamentals and interface contracts | This page exit: know when to change weights instead of prompts or retrieval

SFT (Bridge) ​

Bridge page: this page answers only the engineering-decision question. Training objectives, loss functions and data recipes are derived in Learn LLM (SFT chapter).

What problem it solves ​

Supervised Fine-Tuning (SFT): take an existing pretrained model and continue training it on your input → expected output labeled pairs, making it specialized for a domain, format, or style. It changes model weights — a different road from prompt engineering (changeable per call, zero training cost) and RAG (external knowledge base, model untouched).

What SFT is good at:

  • Domain language: terminology-dense fields such as medicine and law.
  • Stable output format: always producing the same JSON shape or code skeleton, more reliably than prompt constraints.
  • Style and tone: brand writing style, a consistent response personality.
  • Smaller model, lower cost: distill general-model capability into a smaller, faster, cheaper specialized model.

What SFT is bad at (use mainline solutions):

  • Frequently changing knowledge (catalogs, news) → RAG; update the index, take effect immediately.
  • One-off format constraints → structured output + prompt; live in minutes.
  • Prototype stage → any training investment here is premature optimization.

When prompt and RAG are not enough ​

Put SFT on the agenda only when all of the following hold (any miss sends you back to the mainline):

  1. Prompt and few-shot exhausted: format instability and style drift cannot be fixed at the prompt layer.
  2. RAG optimized: hybrid retrieval, reranking, and chunk tuning done; accuracy still short, and the bottleneck is confirmed as "the model cannot express this domain" rather than "retrieval misses".
  3. A meaningful volume of high-quality labeled data: hundreds to tens of thousands of real input/output pairs (order of magnitude, unverified — thresholds vary enormously by task and model; run small experiments first).
  4. Budget and ML expertise: training, evaluation, and regression testing form a continuous pipeline, not a one-off action.

Decision impact table ​

DimensionPromptRAGSFT
Data requirementNone (write the prompt)Unstructured docs sufficeHigh-quality labeled pairs at scale (unverified)
Cost structureInference onlyInference + retrieval infraTraining compute + labeling + eval pipeline (vendor pricing applies; this repo cites no specific figures)
Time to effectImmediateImmediate (re-index)Requires retraining to update
Inference latencyUnchangedAdds retrieval overheadNo retrieval overhead after training; a smaller model can cut latency
Maintenance burdenLowMedium (data freshness)High (data versions, model versions, regression evals)
DebuggabilityHigh (read the prompt)High (inspect retrieval hits)Low (weights are unreadable; eval sets speak)

Managed fine-tuning services ​

If SFT is confirmed, prefer managed services over a self-built GPU pipeline: OpenAI Fine-Tuning, Azure OpenAI, Together AI, Anyscale and others offer fine-tuning APIs. Note: even with managed services, data preparation, effect evaluation, and knowing when to stop training still require ML judgment; an app engineer's role is to collaborate with ML engineers and consume the trained artifact through APIs.

Pre-upgrade checklist ​

Before proposing "let's do SFT", tick every item; any unticked item sends you back to its layer first:

  • [ ] Format problems already attacked with schema validation + few-shot (L1 structured output)
  • [ ] Knowledge problems already attacked with hybrid retrieval + reranking + chunk tuning (L3 advanced retrieval)
  • [ ] At least two or three models compared, ruling out the cheaper explanation of "wrong model chosen" (L5 cost & performance)
  • [ ] Labeled data actually exists with inspectable quality — not "theoretically we could have people label"
  • [ ] A rollback path exists if post-training quality drops (back to the prompt + RAG version)
  • [ ] An ML engineer or managed service owns training and evaluation — not a frontend engineer self-teaching on production

Common anti-patterns ​

  • "Fine-tune it and see": treating the most expensive lever as the experimental one. The correct order is cheapest first — each higher layer's debugging cost is an order of magnitude lower.
  • Chasing knowledge with SFT: pouring catalogs and FAQs into the training set. Knowledge changes force retraining; this is the textbook RAG scenario.
  • Vibes over evaluation: "feels smoother after training" is not evidence. Pre/post comparison must run through eval sets (Evaluation (bridge)).
  • Padding data with generation: generating training data with the model itself amplifies existing biases; human verification cannot be skipped.

Deep derivations and implementation ​

  • Training objectives, data recipes, hyperparameter choices → the SFT chapter of Learn LLM.
  • Mainline next steps: return to RAG and Evaluation (bridge) — "SFT works better" must be backed by evaluation evidence, not vibes.

Built for frontend engineers · Powered by VitePress