Skip to content

Group: Advanced (bridge) | Previous group exit: run long-term with release gates and versioning | This page exit: know which engineering decision this topic affects, and when to go to Learn LLM

Interpretability (Bridge) ​

Bridge page: mechanistic interpretability — superposition, sparse autoencoders, feature decomposition, attribution graphs — with derivations and experiments lives in Learn LLM (Chapter 7 · Transformer internals). This repo answers only the app engineer's question: which of my decisions it changes, and when what you need is actually something else.

1. Overview ​

What problem it solves: the model is a black box — you observe inputs and outputs, but "why it said that" hides in the middle layers. Interpretability research opens those layers: decomposing activations into readable "features", then tracing causal chains between features (attribution graphs), turning "why" into inspectable evidence.

Why app engineers need positional sense: you are not going to dissect the model; you are managing three engineering responsibilities created by the black box:

  • Trust boundary: how deep a behavioral guarantee can you make externally? "We cannot explain why it refuses or doesn't" is a compliance gap in high-stakes settings, not an academic one.
  • Debugging expectations: application-layer tools (evals, traces, ablations) answer only "did behavior change"; "why inside the model" is beyond the application toolbox — knowing this boundary keeps you from wasting time at the wrong layer.
  • Narrative calibration: "planning" or "hidden goals" shown by attribution graphs are products of research tooling; treat them as an evidence framework in communication, not proof that "the model is conscious".

2. Usage ​

The decisions it shapes (positions, not code):

DecisionThe positional sense interpretability gives
Model selection for high-stakes flowsBlack-box depth is a selection cost: the less explainable the behavior's origin, the more external guardrails (eval gates, human review) must compensate
Audit promises to customers/regulatorsPromise only what the application layer can prove (traces, eval records); "internally explainable" is currently a research frontier, not a deliverable
Debugging "why this output"Application layer first: trace and context reconciliation (→ Observability), input ablation, eval-set comparison; only when all point to "that's just the model" does this section apply
Product copy and mental models"Features" and "circuits" are research vocabulary; speak in behavioral contracts (schemas, acceptance) in product contexts

3. Principles ​

(Deliberately minimal: derivations and experiments belong to Learn LLM Chapter 7.)

One-line chain: neurons encode multiple concepts at once (superposition) → dictionary learning / sparse autoencoders decompose activations into monosemantic "features" → causal tracing between features yields attribution graphs. This is a research line, not an application API — app engineers have no callable explain().

4. Development ​

Symptom → where to go:

  • "Output is unstable" → not this section; go to 03-context (contracts that narrow the degrees of freedom).
  • "Why was this answer wrong" → first Observability (trace reconstructs the context) and Evaluation (behavioral boundaries); most "why" turns out to be a context problem.
  • "We must prove model safety to a regulator" → the evidence chain in 08-production; interpretability as research background, not as a promise.
  • Genuinely want to read the internals → Learn LLM Chapter 7 (Transformer internals).

5. Resource Library ​

Pages in this section ​

PageTopic
Reasoning Models & TTCbudgets, latency, and stop conditions for thinking-style output
MoE & Frontier Architecturesthe architectural explanations behind price/speed/context caps
Multimodalminimal usage and boundaries of vision and other cross-modal abilities

Primary sources (as published by the researchers; no content re-narrated):

NameOriginIdentifier
Toy Models of SuperpositionAnthropic (Transformer Circuits)arXiv:2209.10652
Sparse Autoencoders Find Highly Interpretable Model DirectionsCunningham et al.arXiv:2309.08600
Towards MonosemanticityAnthropic (Transformer Circuits)https://transformer-circuits.pub/2023/monosemantic-features/index.html
On the Biology of a Large Language Model (attribution graphs)Anthropic (Transformer Circuits, 2025-03)https://transformer-circuits.pub/2025/attribution-graphs/biology.html

(retrievedAt 2026-09-01; follow Transformer Circuits and lab publications for progress — this repo maintains no snapshots.)

Where learn-ai stops / where to go next ​

  • Feature decomposition, attribution-graph methods and experiments: Learn LLM Chapter 7 (Transformer internals).
  • The "why" tools that actually land for applications: traces and evals (08-production).

Built for frontend engineers · Powered by VitePress