Skip to content

Group: Map · Orientation and Boundaries | Previous group exit: none | This group exit: you can pick the lowest-complexity solution for any requirement and state its upgrade trigger Prerequisites: Tech Map | Next: Site Boundaries and Knowledge Ownership; to enter the Context group now, start with Prompt Engineering

1. Overview ​

Lead with the answer: for any AI requirement, find your rung on the ladder below, adopt that rung's lowest-complexity solution, and climb only when an explicit upgrade trigger fires. This table precedes any protocol introduction — MCP, A2A, ACP, and AG-UI are candidates for rung 6, not a learning order.

Mental model: one ladder ​

The ladder has one rule: default to the lowest rung and let evidence push you up. Every climb adds state, boundaries, and failure modes; every descent removes a class of trouble you have not met yet.

The six-rung decision table ​

RungRequirementLowest-complexity solutionUpgrade triggerDowngrade counterexample
1Fixed input → fixed outputPrompt + structured outputExternal facts or actions neededIf the answer fits in the prompt / context, do not build a retrieval pipeline
2Needs external factsRetrieval / RAG (Retrieval-Augmented Generation)Many sources; permissions / updates / latency become the problemWith a small, stable knowledge set, inlining into context is cheaper
3Needs one controlled actionFunction call / tool callingMulti-step, recoverable, approval-gated → rung 4A single-step action needs no agent loop
4Needs multi-step decisionsWorkflow / same-domain subtasksAutonomous loop with explicit stop conditions → rung 5If steps can be enumerated statically, an autonomous loop only adds risk
5Needs an autonomous loopAgent (runtime + stop condition + permission boundary)Cross-trust-domain, async tasks, capability discovery, cross-implementation interopIf the task completes inside one process, do not cross boundaries with protocols
6Cross-process / org / trust domainProtocol by connection direction (MCP / A2A / ACP / AG-UI)— (top of ladder)In-host tool access: a direct function or plain HTTP suffices

Decision comparison (what each rung adds) ​

RungDirectionControlAdded stateTrust domainLowest complexity
1ReadYou write the prompt; changeable per callnonein-processprompt + output schema
2ReadYou control indexing and refreshindex + data freshnessdata-source boundaryretrieval
3Write (single)function boundary + parameter validationcall results and side effectsin-processtool calling
4Write (multi)orchestration; steps enumerable staticallymulti-step state and intermediatesin-processworkflow
5Read-write loopstop condition + permission boundarysession / memory / checkpointswithin hostagent runtime
6Read-write across boundariesprotocol negotiation + capability discoverytasks / messages / artifactscross-trust-domainboundary protocol

"Direction" means read world (retrieval, citations — no side effects by default) versus write world (tools, actions — they change external state). Every write-world rung must explicitly handle permissions, idempotency, cancellation, retry, and rollback.

When to use / when not to ​

  • Use: before implementation, in design reviews, and whenever judging "should we adopt an agent / a protocol".
  • Do not use: for the concrete implementation once the rung is decided — go to the matching layer's five-part chapter.

Historical milestones: this ladder was frozen in 2026-09 with Issue #116, merging the old "training decision tree" and the "workflow vs agent" discussions; earlier origins are unverified and will not be fabricated.

2. Usage ​

This page is a decision tool with no runnable code artifact; "Usage" = decision drills (pen and paper, ≤15 minutes). Acceptance: all four scenarios select the lowest-complexity solution and can say why a higher rung is not justified.

Drill: four scenarios ​

For each scenario, pick a rung yourself first, then compare with the expected output.

Scenario A: classify support emails into three categories with a confidence score.

  • Expected: rung 1. Fixed input → fixed output; prompt + output schema is enough.
  • Counterexample check: no external facts (no ticket lookup) and no actions (no status change), so no upgrade trigger fires.

Scenario B: answers must be grounded in 8,000 internal wiki pages.

  • Expected: rung 2. External facts exceed the context budget; use retrieval / RAG.
  • Counterexample check: with only 20 pages updated quarterly, inlining into context (rung 1 + long context) is cheaper.

Scenario C: after user confirmation, file the expense report into the approval system.

  • Expected: rung 3. One controlled action: tool call + parameter validation + permission boundary.
  • Counterexample check: the model is not deciding a "check limit → write → notify" chain; if it must, that is rung 4.

Scenario D: a research assistant reads literature, cross-checks, and writes a survey, deciding when to stop on its own.

  • Expected: rung 5. An autonomous loop presumes you can write explicit stop conditions (page cap, time cap, verification passed) and permission boundaries (read-only sandbox).
  • Counterexample check: if the steps are really a fixed "retrieve → summarize → aggregate" chain, that is a rung-4 workflow, not autonomy.

Boundaries of use ​

  • The ladder decides the system shape, not model capability: the same model serves every rung.
  • Rungs can compose (an agent internally calls RAG), but the externally visible shape takes the highest rung involved, and the risk budget is assessed at that rung.

Ladder stepper: click a rung for details

NeedFixed input → fixed output
Minimal solutionPrompt + structured output
Upgrade triggerExternal facts or actions needed
Downgrade counterexampleNo retrieval pipeline when the answer fits the prompt
Read →

3. Principles ​

Why default to the lowest rung ​

Each upgrade monotonically increases three kinds of cost:

  1. State: rung 1 is stateless; rung 2 adds an index and freshness; rung 4 adds multi-step intermediates; rung 5 adds sessions, memory, and checkpoints. More state, costlier recovery and debugging.
  2. Boundaries: from in-process to the data-source boundary, then across trust domains. Every boundary adds a class of authorization, versioning, and protocol-negotiation problems.
  3. Failure modes: a rung-1 failure is "output violates schema", retryable on the spot; a rung-5 failure may be "three side effects already executed", requiring compensation and rollback.

Low-rung failures are local and replayable; high-rung failures are distributed and side-effecting. Hence "let evidence push you up": climbing without a fired trigger means pre-paying complexity for problems you do not have.

Read world / write world as a cross-cutting invariant ​

Rungs 1–2 mostly read the world: they supply evidence with no side effects by default. From rung 3 the system writes the world: it changes external state, so permissions, idempotency, cancellation, retry, human approval, and rollback must be handled explicitly. These are not rung-specific topics but questions every rung from 3 upward must answer — layers 4 and 5 return to this checklist repeatedly.

Why protocols sit at the top ​

Protocols (rung 6) solve communication and capability discovery across trust domains. They add not just technical cost but organizational cost: version negotiation, security boundaries, interoperability commitments. Only when there is an independent trust domain, async tasks, capability discovery, or cross-implementation interop is a protocol the right option. In-host tool access needs only a direct function call or plain HTTP.

Specification vs local measurement ​

This page is a decision pattern with no specification to test against. "Spec vs measurement" columns for each rung's solution live in the matching layer chapters: structured output and model API (Inference & Interface), RAG (Grounding), tool execution (Action), protocols (Interoperability).

4. Development ​

This page has no code integration; "Development" = using the ladder as a design and review gate.

Symptom → Evidence → Action → Done ​

Symptom: the system fails randomly, is hard to debug, and nobody can name the components a single request passed through. Evidence: agent / protocol nodes appear in the architecture diagram while the requirement lands on rungs 1–3; traces show many loop steps unrelated to the task. Action: downgrade to the lowest rung that satisfies acceptance; record the removed capabilities as explicit TODOs pending real triggers. Done: the design doc states the current rung, the checked upgrade triggers, and evidence for any trigger that has fired.

Symptom → Evidence → Action → Done ​

Symptom: the team argues over "should we go agent" with no resolution. Evidence: neither side can write the autonomous loop's stop condition or permission boundary — empty fields mean rung 5's admission criteria are unmet. Action: implement as a rung-4 workflow first; collect cases where steps cannot be statically enumerated as upgrade evidence. Done: either the static step list runs (stay at rung 4), or real enumeration-failure cases accumulate before upgrading.

Symptom → Evidence → Action → Done ​

Symptom: a review demands MCP / A2A "for future integration". Evidence: all collaboration today is inside one host / process; there is no second trust domain, no async task, no capability-discovery need. Action: check each rung-6 upgrade trigger; with none fired, use an in-process solution. Done: the design doc states "if X appears later (cross-trust-domain / async / capability discovery), then introduce protocol Y".

Anti-pattern list ​

  • Noun-driven architecture: pick a trending protocol / framework first, then hunt for a scene that justifies it.
  • Level skipping: jump to agents while skipping output schemas and failure acceptance (the rung-1 homework) — every upper-layer fault collapses into "the model is bad".
  • Demo as acceptance: a working demo ≠ existing stop conditions, permission boundaries, and rollback paths. This is the classic "looks successful but under-evidenced" path.

5. Resource Library ​

Four-level reading route:

  • Beginner: this page + the Tech Map decision tree; recite the six rungs and their triggers.
  • Builder: enter the Context group and make rung 1 solid (prompt + schema + failure acceptance).
  • Operator: enter the Agent Systems group and the Production group for permissions, idempotency, and rollback at rungs 3–5.
  • Researcher: read the L1 resources below on the workflow-vs-agent boundary argument.

Resource table ​

NameEvidence levelCanonical URLPurposeSupported claimNext
Building Effective Agents (Anthropic)L1 (maintainer)https://www.anthropic.com/research/building-effective-agentsThe workflow / agent boundary; the case for starting with compositionsThe "compose first, escalate later" engineering position (retrievedAt 2026-09-01, HTTP 200)Compare with rungs 4 / 5 here
MCP official siteL0 (official spec)https://modelcontextprotocol.io/One rung-6 candidate: the agent ↔ tools / data boundaryProtocols solve boundary communication (retrievedAt 2026-09-01, HTTP 200)MCP chapter
A2A official siteL0 (official spec)https://a2a-protocol.org/One rung-6 candidate: cross-trust-domain agent collaborationRemote collaboration protocols presume cross-trust-domain needs (retrievedAt 2026-09-01, HTTP 200)A2A chapter
Learn LLM chapter 11 (RAG)siblinghttps://llm.zenheart.site/chapters/11-ragDeep principles behind rung 2 (when retrieval works)Retrieval internals belong to Learn LLM (retrievedAt 2026-09-01, HTTP 200)This repo's Grounding group

Active falsification and open questions ​

  • Falsification entry: if you find a scenario where rung N's lowest solution cannot pass acceptance while rung N-1's triggers have not fired, the ladder has a hole — revise the triggers instead of hiding the scenario.
  • Open: the rung-6 "choose protocol by connection direction" specifics (which boundary fits MCP / A2A / ACP / AG-UI) land in the layer-4 protocol map; this page keeps only the decision frame.

Where learn-ai stops / where to continue ​

  • Implementation per rung: this repo's group chapters.
  • Deep principles for rungs 2 / 5 (retrieval math, how training shapes behavior): Learn LLM.
  • Methods proving each rung is "done": evals.

Built for frontend engineers · Powered by VitePress