Skip to content

Group: Advanced (bridge) | Previous group exit: run long-term with release gates and versioning | This page exit: know which engineering decision this topic affects, and when to go to Learn LLM

MoE and Frontier Architectures (Bridge) ​

Bridge page: MoE routing, load balancing, MLA attention derivations, and long-context position encodings live in Learn LLM (Chapter 9 · Inference and Quantization; Chapter 20 · DeepSeek topic). This repo keeps only the application-side positional sense: these architectures explain the "why" behind the price, speed, and context caps in your selection table.

1. Overview ​

What problem it solves: three frontier directions each break an old equation:

  • MoE (Mixture-of-Experts) breaks "more parameters = more compute per token": total parameters can be huge while each token activates only a small set of experts — capacity and per-call cost decouple.
  • MLA (Multi-head Latent Attention) breaks "long context = linear KV-cache explosion": compressing KV into a low-dimensional latent space flattens the inference-cost curve for long contexts.
  • Long context breaks "the model can only see a small window at once": but "fits in the window" does not mean "used well" — utilization of information in the middle of long inputs has a known degradation shape.

Why app engineers need positional sense: you do not need to derive MLA, but the answers to "why is this vendor cheap", "why is that one fast", and "how full should I pack the context" all come from these three lines. Without them, selection degenerates into memorizing a price list.

2. Usage ​

DecisionThe positional sense the architecture gives
API selection (price/speed)an MoE model's per-token cost is set by activated parameters, not total parameters — "huge total parameter count" is not a reason it's expensive, "large activated parameters" is; hosted APIs have already folded this into pricing
Hosted vs self-hostedMoE saves compute, not memory: all expert weights must stay resident. The self-hosting gate for MoE is memory and communication, not FLOPs — one source of the "hosted is cheap, self-hosted is expensive" paradox
Context strategylong context is not a free RAG replacement: a full window costs real money and mid-window utilization degrades. The doctrine stands — "what retrieval should do, leave to retrieval" (→ 04-grounding); use long windows for what must be present in full (entire contracts, long code files)
KV-cache-related costscache hits and prefix reuse (prompt caching) vary by architecture; for long-context multi-turn workloads, keeping prefixes stable is an engineering money-saver (→ Cost and Performance)

3. Principles ​

(Deliberately minimal: derivations belong to Learn LLM.)

  • MoE in one line: a router picks a few experts per token — capacity grows, per-call compute doesn't; the price is load balancing and routing stability.
  • MLA in one line: compress KV into a low-dimensional latent representation and restore on demand — an extension of the MQA/GQA compression line.
  • Long context in one line: windows can grow, but mid-window utilization degrades (the lost-in-the-middle phenomenon) — where you place information inside the context remains an engineering variable.

Deep water: Learn LLM Chapter 9 (KV cache, MQA/GQA), Chapter 20 (DeepSeek topic: MLA/MoE/GRPO).

4. Development ​

Symptom → where to go:

  • "Switched to a supposedly stronger model and the bill changed" → recompute unit cost through the activated-parameter lens (→ Cost and Performance).
  • "Do we still need RAG with million-token models" → yes. Grounding for private/fresh facts does not vanish with bigger windows (→ 04-grounding); long windows change what is feasible for "present in full" scenarios.
  • "Key information in the middle of long documents keeps getting missed" → placement strategy: critical facts at the head or tail, or structured chunking plus retrieval — do not count on uniform attention across a full window.
  • Want the routing/MLA derivations → Learn LLM Chapters 9 and 20.

5. Resource Library ​

Primary sources (identifiers only; no content re-narrated):

NameOriginIdentifier
Outrageously Large Neural Networks (sparse MoE origin)Shazeer et al.arXiv:1701.06538
Switch Transformers (Top-1 routing)Fedus / Zoph / ShazeerarXiv:2101.03961
Mixtral of Experts (open MoE representative)Mistral AIarXiv:2401.04088
DeepSeek-V2 (introduces MLA)DeepSeek-AIarXiv:2405.04434
DeepSeek-V3 Technical ReportDeepSeek-AIarXiv:2412.19437
Lost in the Middle (long-context position effects)Liu et al.arXiv:2307.03172

(retrievedAt 2026-09-01; current architectural details of each model follow its technical report and official docs.)

Where learn-ai stops / where to go next ​

  • MoE/MLA/GRPO derivations and toy reproductions: Learn LLM Chapter 20 (DeepSeek topic), prerequisite Chapter 9 (Inference and Quantization).
  • Context-engineering landing points: 03-context; cost closure: 08-production.

Built for frontend engineers · Powered by VitePress