Skip to content

Group: Production | Previous group exit: can restrict permissions, pause/resume tasks | This group exit: can substantiate quality, risk, performance, rollback, and ownership with replayable evidence Prerequisites: Recovery and Human-in-the-Loop | Next: Testing, Evaluation (Bridge), Observability, Security, Cost and Performance, Deployment and Release; after this group, the Resource Library

1. Overview ​

BLUF: Between "the feature works" and "we can run this in production" lies an entire group's worth of work. A demo needs one success; operations need replayable evidence — the ability to reconstruct what happened during an incident, to predict the impact of a change, and to point at a rollback path and an owner. This group consolidates the engineering of the upstream groups into five kinds of launch evidence: quality, risk, performance, rollback, and ownership.

Mental model: six evidence chains ​

The six chains are not six sequential stages; they are six facets of the same artifacts. One trace is simultaneously observability evidence, cost evidence, and incident-replay material.

Symptom routing table ​

Symptom ("the feature works, but…")Go toEvidence you gain
Change one prompt, no idea what brokeTestingDeterministic test pyramid: assertable blast radius
"Better/worse" is only a vibeEvaluation (Bridge)Golden set + threshold as a release gate (methodology → evals site)
Production incident, only scattered logs to guess fromObservabilityReplayable trace/span + correlation ID
Worried about injection, overreach, leaks, but "the model will refuse"SecurityCombined defense: input filter + permission boundary + output scanner
Nobody can explain the monthly bill or P95 latencyCost and PerformancePer-request cost tracking + latency breakdown + optimization ladder
Releases by manual copy, rollback by memoryDeployment and ReleaseGate chain, versioned rollback, runbook

Decision table: how each evidence type is obtained ​

Evidence typeDirectionControlStateTrust domainMinimum complexity
Quality (testing + eval)ReadYou write assertions/datasetsFixtures + golden setInside CIOne node:test file
Risk (security)Read/WriteFilters + permission boundary + auditRules + audit logApplication trust domainOne layer of I/O filtering
Performance (cost + latency)ReadSampling & budget configTime seriesObservability backendPer-request usage accounting
Rollback (deployment)WriteVersion + traffic switchVersioned artifactsDeployment trust domainVersion number + dual-write switch
Ownership (runbook)—Roster + escalation pathDocs + alert routingOrganizationOne page of who owns what

Cross-cutting statement ​

Security, cost, and observability do not first appear in this group: permission boundaries are introduced in Tool Execution, token usage arrives with the usage field of the Model API, redaction matters as soon as Context Engineering begins. The upstream groups handle these concerns locally in their own scenarios; the Production group consolidates them into systematic launch evidence — no longer "done right somewhere", but "provable everywhere".

When to use / when not to ​

  • Use: any system crossing from "personal demo" to "has users, has on-call, has a bill".
  • Do not use: one-off scripts, local experiments, learning toys — building five evidence chains for them is over-engineering. Judge with the Complexity Decision Ladder.

Historical milestone: this group's structure was frozen in 2026-09 with Issue #116, upgrading and merging the legacy engineering/ pages (testing/evals/observability/security/cost-optimization) plus the testing/ and evaluation/ directories; earlier provenance unverified, not fabricated.

2. Usage ​

This page is the routing layer and has no runnable artifact; "usage" = a launch-readiness self-audit (on paper, ≤15 min). Acceptance: for the system you want to launch, every one of the five evidence types resolves to "concrete artifact + named owner" — any blank cell is your gap.

Drill: the five-question audit ​

Answer each question for the target system; record "where the evidence lives (file/dashboard/doc)" and "who owns it":

  1. Quality: on the last prompt change, which tests ran in CI? How much did the eval score move? (→ Testing / Evaluation)
  2. Risk: an injection payload hits your input endpoint — which layer blocks it, and where is the block logged? (→ Security)
  3. Performance: pick a random production request from yesterday — can you find its token cost and per-segment latency within a minute? (→ Observability / Cost and Performance)
  4. Rollback: what is the previous version of the current one? Which exact actions switch back to it? (→ Deployment and Release)
  5. Ownership: at 2 a.m. the model vendor returns 5xx — which dashboard does the on-call look at first, and what is the first action? (→ the runbook in Deployment and Release)

Usage boundary ​

  • This group does not teach "how to test / how to instrument" — that is the six sub-chapters' job; this page only routes symptoms to the right evidence chain.
  • The depth of each evidence type scales with business risk: an internal tool does not need bank-grade audit logs, but it does need the minimum version that answers the five questions.

3. Principles ​

Why consolidation happens in the Production group ​

Every upstream group exits on "it works": the schema validates, the interaction is cancellable end to end, retrieval is traceable, permissions pause and resume. But "works" is a single-run property; operations demand a continuing property — the ten-thousandth request as reliable as the first, the tenth change as predictable as the first. The instrument that upgrades single-run to continuing is evidence:

  • Predictable behavior ← testing + evaluation (comparable before/after change)
  • Explainable failures ← observability (replayable traces)
  • Verifiable boundaries ← security (attack paths with block records)
  • Accountable spend ← cost & performance (per-request accounting)
  • Reversible changes ← deployment (versions that switch back)

Replayability is the group-wide invariant ​

All five evidence types share one acceptance criterion: it can be replayed afterwards. Tests can re-run, datasets can reproduce, traces can be exported and re-inspected, attack samples can be re-fired against the filter, costs can be recomputed per request, versions can be redeployed. Any observation of the form "we saw it then, can't find it now" is not evidence — that is the line between real observability and "happened to see it".

Division of labor with external owners ​

This group deliberately stops at two places: evaluation methodology (datasets/scorers/judges/statistics/red-teaming) is owned by evals — this repo only answers "when is evidence needed and how it plugs into the release gate"; the mechanics of the model-alignment risk layer (reward hacking, jailbreak propensity) are owned by Learn LLM — this repo only does application-side engineering. Details in _phase0/bridge-register.md.

Spec vs local measurement ​

This routing page has no spec to implement; each sub-chapter carries its own "spec vs measured" column (e.g. OTel GenAI conventions, the OWASP 2026 list, Vitest version requirements).

4. Development ​

This page has no code integration; "development" = using the five-question audit as a launch-review gate.

Symptom → Evidence → Action → Done when ​

Symptom: review gets rejected with "how do you prove it's correct?" and the team can only replay the demo. Evidence: no five-question audit in the review materials; no reconcilable artifacts between change history and production behavior. Action: fill in the audit; for any blank cell, build the minimum version first (one test file, one trace, one page of runbook all count). Done when: the review materials contain the audit table with concrete artifact paths and named owners per row.

Symptom → Evidence → Action → Done when ​

Symptom: the incident postmortem runs three pages and concludes "the model misbehaved". Evidence: no exportable trace in the incident window; request logs carry only status codes, no token usage or tool-call records. Action: add correlation-ID propagation and span recording per Observability; put "replay the incident window" first in the postmortem template. Done when: the next similar incident starts its postmortem by exporting the window's traces instead of scrolling chat logs.

Symptom → Evidence → Action → Done when ​

Symptom: the bill doubles and nobody knows which feature, user, or request class caused it. Evidence: the only cost number is the vendor's monthly invoice; no per-request/per-feature usage accounting. Action: establish per-request cost tracking per Cost and Performance (usage field → record → aggregate); decomposition before optimization. Done when: the month's cost splits along "feature × model × input/output/cache" and you can name the largest item.

Anti-patterns ​

  • Evidence theater: screenshots and numbers assembled to pass review, not replayable afterwards — the Production-group form of "looks successful but under-evidenced".
  • All six chains, fully armed: bank-grade audit and full-rate trace sampling for an internal tool turns evidence cost into the new operational burden.
  • Evidence detached from the system: tests run against mocks, traces enabled only in staging — the production path is exactly the uncovered one.

5. Resource Library ​

Four-level reading route:

  • Beginner: read this page's routing table; run a five-question audit on your own system; find the largest gap.
  • Builder: enter Testing and Observability — the two chains that compound first (change safety + explainable failures).
  • Operator: enter Security, Cost and Performance, Deployment and Release; fold risk, spend, and rollback into daily routine.
  • Researcher: read the L0/L1 specs below (OWASP, NIST, OTel) for the field's framing of "reliable AI operations".

Resource table ​

NameEvidence levelCanonical URLPurposeSupported claimNext
OWASP GenAI LLM Top 10 2026L0 (official list)https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/Baseline for risk evidencePublished 2026-08-04; ranking calibrated on 7,714 real incidents (retrievedAt 2026-09-01)Security
NIST AI RMF 1.0L0 (official framework)https://www.nist.gov/itl/ai-risk-management-frameworkVoluntary governance frameworkReleased 2023-01-26, voluntary; GenAI Profile (NIST AI 600-1) 2024-07-26; NIST notes 1.0 is being revised under the White House AI Action Plan (retrievedAt 2026-09-01)Security governance section
OpenTelemetry GenAI semantic conventionsL0 (official spec)https://github.com/open-telemetry/semantic-conventions-genaiNaming baseline for observability evidenceGenAI conventions moved to a dedicated repo in 2026-06; covers GenAI clients/MCP/provider conventions (retrievedAt 2026-09-01)Observability
evals sitesibling (cross-repo owner)https://evals.zenheart.site/Sole owner of evaluation methodologyThis repo keeps only release-gate integration (bridge-register 2026-09-01)Evaluation (Bridge)

Falsification and open questions ​

  • Falsification entry: if you find a real system that operates reliably while lacking one of the five evidence types, the necessity assumption for that type has a hole — first check whether it uses an equivalent evidence form outside this table, then consider revising this group's exit.
  • Open: how to grade the depth of the six chains by business risk (which scenarios may explicitly waive which evidence) — to be backfilled as sub-chapter runbooks accumulate cases.

learn-ai stops here / where to go next ​

  • Evaluation methodology, dataset construction, judges and statistics: evals.
  • How model alignment and post-training shape risk behavior: Learn LLM.
  • The full resource index after this group: Resource Library.

Built for frontend engineers · Powered by VitePress