Skip to content

Group: Context (group guide) | Previous group exit: interface contracts and serving shapes | This group exit: assemble a single-turn context within budget and manage cross-turn memory

Context Group: What the Model Sees This Turn ​

Group: Context group | Before this group: can write and validate input/output schemas (structured-output) | Group exit: the full chain of "what the model sees this turn" — a budget you can compute, sources you can assemble, rot you can counter Prerequisites: structured-output | Next: Embeddings and Retrieval, Tool Calling Contract

1. Overview ​

Bottom line: however well the prompt is written, it controls only "how to say"; the other half of quality and cost lies in "what the model sees". The Context group splits that question into five orthogonal sub-questions, one page each: how to express (prompt), how much fits (context-window), what to load (context-engineering), where cross-turn history lives (session-memory), how to declare repo conventions (repo-context).

Why a whole group: context is a finite resource with diminishing returns — the more tokens in the window, the worse the model's accurate recall (context rot, present across models); meanwhile every resent token is billed. Do input-side curation badly and every layer above (grounded retrieval, tool execution, operations reconciliation) pays for noise.

Mental model: the transformation chain from intent to window ​

text
human intent
   │ ① prompt              how to express (task/constraints/examples/output format)
   ▼
controlled instruction ── ② context-window      how much fits (token budget, pairs, output reserve)
   │
   │ ③ context-engineering  what to load (sources, priorities, compaction, staleness)
   ▼
assembled window ── ④ session-memory     where history lives (storage, trimming, recovery, concurrency)
   │
   ▼ ⑤ repo-context        how conventions are declared (AGENTS.md, nearest-wins, host injection)
what the model actually sees this turn

Symptom routing: which page takes which problem ​

text
Unstable answers, off-target replies (expression)          → prompt
400 over-window, linear cost growth, truncated output      → context-window
Retrieval stuffed in yet still wrong, stale citations      → context-engineering
History lost on refresh, growing amnesia, concurrent bugs  → session-memory
Wrong deps / test commands after switching assistants      → repo-context

Topic navigation table ​

TopicWhat question it answersExitLink
Prompt EngineeringHow to turn intent into an executable instruction?Write four-element prompts managed as codeprompt
Context WindowHow much fits this turn?Compute four-block budgets, spot overflow, trim with pairs intactcontext-window
Context EngineeringWhat should be loaded this turn?Design multi-source assembly: priorities, visible eviction, compaction and stalenesscontext-engineering
Session and StateWhere does cross-turn history live?Build sessions: budget trimming, persistent recovery, concurrency guardssession-memory
Repo ContextHow are repo conventions declared?Land an AGENTS.md with real commands and correct layeringrepo-context

When to enter this group / when not ​

AudienceFrontend / full-stack engineers wiring models into products or workflows who start caring about quality and cost
EnterThe prompt is clear, but any of these appear: multiple turns, external content (files / retrieval / tool results), coding assistants
Do not enterAnswers still unstable or unparseable — close the contract loop first at structured-output; model internals → Learn LLM
Not this groupBuilding the retrieval index → grounding group; executing actions safely → tool-calling

Decision table: how the five topics divide the work ​

promptcontext-windowcontext-engineeringsession-memoryrepo-context
ControlsExpression of instructionsCapacity of inputContent selection of inputCross-turn history I/ORepo-level conventions
DirectionHuman → model (intent)System → window (budget)Sources → window (curation)Session → store → windowRepo → host → window
ControlText, fully yoursAssembly layer, fully yoursAssembly layer, fully yoursStore implementation, yoursRepo file + host reading
StateVersioned textRecomputed per turnReassembled per turnPersistent across requestsEvolves with the repo
Trust domainDiffableEstimate vs exact countingSource freshnessStorage and concurrencyCommand truthfulness
Min complexityLowest, always try firstLow (single turn) to medium (multi-turn)MediumMedium (+ persistence)Low (one text file)

Read in navOrder: prompt → context-window → context-engineering → session-memory → repo-context.

Historical milestones: unverified (vendor-capability timelines live on each topic page; this page repeats none).

2. Usage ​

This page is navigation and has no standalone fixture — the group's unified hands-on exit is context-window's zero-key budget allocator (15 minutes: four-block planning → overflow alarm → pair-preserving trim, negative case included). It is the group's meeting point: the prompt occupies the system block, history the history block, retrieval the retrieval block, output the output reserve — five topics converge on one budget sheet.

15-minute self-check (after running the fixture):

  1. Why must the output reserve count toward the budget? (A: input and output share the window; without a reserve, stop_reason: max_tokens truncates)
  2. Which two kinds of messages must trimming never sever? (A: the system head; tool_use / tool_result pairs)
  3. What is the first action when over budget? (A: evict the lowest-priority source visibly — not silently, and not by switching to a bigger-window model)

Answer all three and the front half of the group's exit is met; then run context-engineering's assembler to add "sources have priorities, eviction is visible".

Acceptance command (the deterministic acceptance of context-window):

bash
npx tsx@4 context-window.ts && echo BUDGET-OK

Cleanup: delete temporary files.

3. Principles ​

Why "context" deserves its own group ​

The naive view of a model API is "prompt in, text out" — apparently one knob: the prompt. Once engineered, "what is seen this turn" splits into at least five sub-problems, each with its own invariants: tokens have budget invariants, sources have priorities and invalidation conditions, history has storage and concurrency, repo conventions have nearest-wins and command truthfulness. Each sub-problem is independently acceptance-testable — that is the reason to split them.

Group-level invariants (expanded per page; the synopsis here) ​

  1. Count before sending: any assembled request can be judged for overflow before it ships (→ context-window I1).
  2. Trimming must not break structure: system kept, tool pairs intact, eviction visible (→ I2 and context-engineering).
  3. Every source has a priority and an invalidation condition: assembly without priorities is arrival-order by another name (→ context-engineering).
  4. Commands must be real: assistants execute what AGENTS.md lists (→ repo-context).

Two universal "stop here" lines ​

  • Attention math, KV cache, the four memory patterns → Learn LLM Chapter 15 (this group takes only the engineering implications).
  • Vendor window numbers, caching prices, token-counting endpoint fields → each vendor's docs on the day (this group maintains no number lists).

4. Development ​

Group exit criteria ​

  • [ ] Rewrite a vague request into a four-element prompt that enters version control (prompt)
  • [ ] Compute a four-block token budget for a request; trim without severing tool pairs (context-window)
  • [ ] Assign priorities to multiple sources with reported eviction and stable sources in the prefix (context-engineering)
  • [ ] Build sessions for multi-turn chat: budget trimming, persistent recovery, concurrency guards (session-memory)
  • [ ] Land an AGENTS.md with real commands and correct layering for your repo (repo-context)

Debug runbooks (group-level symptoms) ​

R1 "Routed to the wrong layer; editing prompts instead of trimming the window" ​

Symptom: multi-turn quality drops; the team ships three prompt revisions with no improvement. Evidence: the per-turn token curve still rises monotonically — the problem is not expression but budget. Action: return to context-window R1 per the symptom routing table; expression symptoms (instability, off-target) belong in prompt. Done when: failure samples are attributed (expression / budget / content / history / conventions) before action; the attribution is recorded in the PR.

R2 The "gets dumber as it goes" localization chain ​

Symptom: answer quality clearly declines in the second half of long sessions. Evidence: check in order — is the token curve near the window (budget) → are retrieval chunks stale (content) → were key turns trimmed (history). Action: budget full → trim or compact; content stale → switch to JIT; history lost → adjust retention (see context-window, context-engineering, session-memory respectively). Done when: similar sessions stop degrading monotonically with turns; the chain is written into the team's troubleshooting doc.

R3 "Behavior changes with every assistant" ​

Symptom: correct in Cursor, wrong deps in Claude Code; every newcomer steps on the same rake. Evidence: no AGENTS.md at the repo root, or commands that do not match reality. Action: land the file per repo-context R1; monorepo exceptions use nested files. Done when: fresh sessions use the right commands on turn one; behavior is consistent across tools and teammates.

Anti-patterns (group level) ​

  • Skipping the budget and piling on retrieval: more content, worse quality — rot does not vanish with a bigger window.
  • Fixing content problems with longer prompts: trim what should be trimmed; change sources that should be changed.
  • Assembly logic scattered in many places: concatenation that bypasses priorities and guards is a second source of truth.
  • Failures swallowed silently: eviction, trimming, and expiry must all be reported, or none of it is debuggable.

5. Resource Library ​

Four-level reading route ​

LevelReadWhy this order
BeginnerAnthropic prompt engineering overviewOpenAI guide, context window section
BuilderThe group's five pages in order + the context-window fixtureA regression-checkable budget and assembly loop in hand
OperatorBoth providers' prompt caching and token counting docsagents.md
ResearcherAnthropic: Effective context engineeringChroma: Context Rot

Resource table ​

NameLevelcanonical URLUseSupportsNext
The group's five topic pagesL1prompt · context-window · context-engineering · session-memory · repo-contextMain path—Read in order
Anthropic context engineeringL1https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agentsMental model overviewMeta-principle / JIT / compaction / rotRead the original
OpenAI prompt engineeringL1https://developers.openai.com/api/docs/guides/prompt-engineeringOfficial budget viewWindows measured in tokensRead their Structured Outputs
Anthropic Context windowsL1https://platform.claude.com/docs/en/build-with-claude/context-windowsWindow and billing basisInput and output share the budgetPair with token counting
agents.mdL1https://agents.md/Repo-level context conventionFormat, nearest-wins, tool supportLand one for your repo
Chroma Context RotL4https://research.trychroma.com/context-rotEmpirical decayLong-context retrieval degradationRead before designing long-document tasks
Learn LLM Chapter 15L2https://llm.zenheart.site/chapters/15-prompt-memoryMechanism bridgeAttention / KV cache / four memory patternsWhen you need the "why"

(retrievedAt: 2026-09-01.)

Active falsification and open questions ​

  • This group's conclusions rest on the two vendors' docs and Anthropic's engineering article as retrieved 2026-09-01; window numbers and caching prices drift — recheck within ≤ 6 months.
  • The five-topic split is this repo's pedagogical organization, serving each page's acceptance-testable exit; it is not an industry-standard taxonomy.
  • Open: no general formula for the context-rot knee (→ context-window); no public benchmark for compaction retention (→ context-engineering).

Where learn-ai stops / where to continue ​

Built for frontend engineers · Powered by VitePress