Skip to content

Group: 01 · Model Lifecycle (bridge group) | Exit of the group above: you can decide knowledge ownership and pick the lowest-complexity option for a need (00-map group) | Exit of this page: you can translate vendor-launch architecture terms (long context, MoE, cache pricing) into engineering variables of context length, inference cost, and input modality Prerequisites: The LLM Mental Model (Bridge) | Next: Data, Pretraining, and Scaling (Bridge)

1. Overview ​

Lead with the answer: application engineers do not need to derive Attention, but they do need position awareness — architecture is frozen at pretraining and directly determines context length, inference cost, and the multimodal path, all three of which show up on your API bill and in your selection spreadsheet. This page gives each component a one-line "what it is / what decision it touches" position; derivations and from-scratch implementations belong to Learn LLM chapters 7, 9, and 20.

One-line positioning table ​

ComponentWhat it is (one line)What decision it touches
TransformerThe self-attention-based sequence backbone; the common foundation of modern LLMsSwitching models = swapping a set of weights with an unchanged interface; far cheaper than switching paradigms
AttentionThe mechanism where tokens weight each other pairwise; the causal mask enables token-by-token generationAttention computation grows with sequence length — the reason long context costs more
Positional encoding (RoPE)Rotates position information into queries/keys so the model senses token orderThe source of a model's advertised context length; window extension is a recurring release variable
MoE (Mixture of Experts)Activates only a fraction of expert parameters per layer: large total, small per-token activationParameter count decouples from per-token cost; an architectural root of vendor pricing differences
KV cache and MQA / GQACaches attention keys/values at inference, plus variants that shrink that cacheThe prefill/decode cost structure and prompt-prefix cache pricing

Why position awareness, not derivation ability ​

Architecture is where the left chain freezes physical constraints into interface properties: after training ends, nobody can change this model's window assumptions or expert routing — you wait for the vendor's next version. Architectural knowledge therefore takes the engineering form of "decoding ability when reading release notes and pricing pages", not write ability. A vendor's "longer context" or "cheaper tier" is almost always an architecture variable — but the landing constraints are whatever the vendor documents say.

2. When to Hop ​

This page has no runnable artifact; route via this table:

Symptom / questionWhere to goWhy
Want a hand-written Attention, why the causal mask existsLearn LLM chapter 7The canonical from-scratch Transformer
KV cache, quantization, prefill/decode cost modelsLearn LLM chapter 9Engineering detail of inference optimization
MLA / MoE accounting / modern architecture topicsLearn LLM chapter 20DeepSeek-family architecture special topic
How API billing and prefix-cache discounts workThis repo's Model API contractApplication-side consumption lives here
How much context a long session should carryThis repo's 03-context groupContext trade-offs are this repo's main line
Impact of a vendor's new architecture version on productionThis repo's 08-production groupVersion-upgrade regression is an operations problem

3. Principles ​

One-paragraph positioning: Transformer replaced recurrence with self-attention for training parallelism, at the price of attention cost growing with length; RoPE determines how far the model can extrapolate context; MoE decouples "total parameters" from "per-call computation" via sparse activation; KV cache turns decode from recomputation into lookup. Those four sentences are all the architecture an application engineer needs — why scaled dot-product divides by √d, the complex-number form of RoPE, and expert-routing load balancing are derived step by step in Learn LLM chapters 7, 9, and 20 (retrievedAt 2026-09-01); this repo does not copy them.

4. Engineering Decision Impact ​

Engineering decisionArchitectural factEngineering action
Long-session cost estimationAttention and KV cache volume grow with lengthSummarize and prune long history; watch vendor prefix-cache pricing
Model family selectionDense and MoE have different cost structuresTier by task; compare actual per-task cost, not parameter count
Context window upgradeWindow extension often changes behaviorRun regression evaluation before switching
Multimodal intakeImages enter token space via projection and are billed per tokenBudget in tokens, not pixels or resolution
Local / edge deploymentQuantized variants shift the accuracy-cost balanceThe math belongs to Learn LLM chapter 9; the selection decision to this repo's 02-inference-interface group

Maintenance rule (in place of runbooks): when Learn LLM chapters 7/9/20 change structure, re-verify this page's deep links and update lastVerified; this page does not maintain a per-vendor adoption matrix of architectures — that drifts with versions; defer to each vendor's model documentation.

5. Resource Library ​

Four-level reading route:

LevelWhat to readWhy this order
BeginnerThis page's positioning table + the group guideBuild the "component → interface property" map first
BuilderThis repo's Model API contractConsume these architectural properties at the interface layer
OperatorCost and performanceBring architecture variables into cost governance
ResearcherLearn LLM chapters 7, 9, 20 + the three original papersDescend to mechanism and derivation

Resource table ​

NameLevelCanonical URLSupported claimNext
Learn LLM ch. 7 · TransformerEhttps://llm.zenheart.site/chapters/07-attentionFrom-scratch Attention / causal mask / multi-headHand-write a Transformer block
Learn LLM ch. 9 · Inference and quantizationEhttps://llm.zenheart.site/chapters/09-inference-cacheEngineering detail of RoPE, KV cache, MQA/GQA, quantizationUnderstand the inference cost structure
Learn LLM ch. 20 · DeepSeek special topicEhttps://llm.zenheart.site/chapters/20-deepseekMLA / MoE accounting and modern architecture evolutionOptional extension reading
Attention Is All You Need (Vaswani et al., 2017)L4https://arxiv.org/abs/1706.03762The original Transformer paperRead the attention section
RoFormer (Su et al., 2021)L4https://arxiv.org/abs/2104.09864Where RoPE positional encoding was introducedRead the rotary section
Switch Transformers (Fedus et al., 2021)L4https://arxiv.org/abs/2101.03961The representative sparse-MoE workRead the expert-routing section

(Learn LLM chapters verified via its chapter index; arXiv links are stable abs URLs; retrievedAt 2026-09-01.)

Active falsification and open questions ​

  • Falsification entry: if a decision truly requires architectural derivation to get right (e.g. implementing an attention kernel yourself), it belongs to Learn LLM or an inference-engine team, not this page — sink the claim instead of expanding here.
  • Open: vendor adoption of MQA / GQA / MLA and long-context extrapolation drifts quickly with versions; this page keeps no adoption matrix — defer to each vendor's model card and pricing page.

Where learn-ai stops / where to continue ​

Built for frontend engineers · Powered by VitePress