Skip to content

Efficient Serving: Making Tokens Cheaper ​

Group: 2 · Inference & Interface | Previous group exit: can decompose one latency into queue / prefill / decode / network | This group exit: can read where the cached input / cache write / cache read fields on the bill come from, and restructure prompts and sessions into cache-friendly shapes Prerequisites: Inference Fundamentals | Next: Model API Contract, Cost and Performance

1. Overview ​

Inference Fundamentals gave the structure: prefill decides TTFT, decode is memory-bound, and the KV cache eats memory linearly with context. This page answers the most practical engineering step: what makes tokens cheaper on the serving side, and how those mechanisms show up on your bill. Read this mapping and you can explain "why cached input has one price and regular input another" — and do the math when choosing between hosted and self-hosted.

Four levers, one sentence each:

  • Quantization: compress weights from 16-bit to INT8/INT4 and similar low precisions — less memory and memory traffic, lower cost and latency; accuracy is the trade.
  • Paged KV: slice the KV cache into small blocks allocated on demand (the PagedAttention idea in vLLM), eliminating the fragmentation of "pre-allocate max context per request" — more concurrent sequences fit the same memory, throughput rises.
  • Prefix caching: reuse the KV of repeated prefixes (system prompts, tool definitions, few-shot examples) and skip prefill — the physical basis for vendors discounting cached input.
  • Speculative decoding: a small model drafts, the large model verifies in parallel, advancing multiple tokens per step — decode latency drops without changing the output distribution.

When to use / when not to ​

  • Use: multi-turn session costs spiraling, a long fixed system prompt, understanding or negotiating serving costs, estimating the upside of self-hosting.
  • Do not use: the math of quantization (how INT8/INT4 compress, how to measure accuracy loss) → Learn LLM chapter 9; API fields and error semantics → Model API Contract; overall cost governance and budgets → Cost and Performance.

Decision table: the four levers compared ​

LeverActs onWho can do itWhere it pays offYour integration actionCost / boundary
QuantizationWeight precisionVendors already do it; self-hosting picks the levelCost, latency, memoryPick a vendor quantization tier or a smaller model; pick a precision level when self-hostingAccuracy loss needs an eval set
Paged KVKV memory fragmentationInside the serving engineThroughput (cheaper unit price, indirectly)Pick a modern engine when self-hosting; nothing when hostedNo direct single-request latency gain
Prefix cachingRepeated prefix computationVendor side (automatic or explicit); self-hosting can enable itCost, TTFTReorder the prompt to "stable first, dynamic last"Hits need byte-identical prefixes + minimum length + TTL
Speculative decodingDecode serializationA vendor-side switch; configurable in self-hosted enginesLatencyChoose models/endpoints that support itOutput unchanged; not available for every model

Historical milestones ​

  • PagedAttention shipped with vLLM (its docs homepage indexes the "vLLM announcing blog post (intro to PagedAttention)" and the SOSP 2023 paper). retrievedAt 2026-09-01.
  • FlashAttention (arXiv 2205.14135, 2022-05): IO-aware attention kernels, a bottom-layer component of the serving stack. retrievedAt 2026-09-01.
  • Vendor prefix-caching launch timelines are unverified; we do not fabricate them.

2. Usage ​

Minimal hands-on: a zero-key prefix cache cost calculator (≤ 15 minutes) ​

No API key, no dependencies, pure deterministic arithmetic. Simulate a multi-turn chat: the system prompt is the cacheable prefix, and every turn resends the whole history (the API is stateless). The price table mirrors the structure of vendor billing — two input prices, cached and uncached — with fictional round numbers; the structure is the point.

Environment: Node ≥ 23.6 (runs .mts directly). Save as prefix-cache-cost.mts:

ts
// fixture: zero-key, zero-dependency prefix cache cost calculator. Mirrors vendor billing STRUCTURE; numbers are fictional.
interface PriceTable {
  inputPerMTok: number;        // uncached input price (per 1M tokens, fictional units)
  cachedInputPerMTok: number;  // cache-hit input price
}

// Structure mirrors: the "Cached input" column of OpenAI's pricing table (automatic prefix caching, >=1024 tokens)
// and Anthropic's cache_read_input_tokens (explicit cache_control breakpoints)
const PRICE: PriceTable = { inputPerMTok: 3, cachedInputPerMTok: 0.3 }; // hit price = 10% of base, mirroring the cache-read tier

interface TurnBill { turn: number; cacheHit: boolean; uncachedTokens: number; cachedTokens: number; cost: number }
interface SessionBill { turns: TurnBill[]; totalCost: number; cacheHitRatio: number }

function costOf(uncached: number, cached: number, p: PriceTable): number {
  return (uncached / 1e6) * p.inputPerMTok + (cached / 1e6) * p.cachedInputPerMTok;
}

// A multi-turn chat: the system prompt prefix is cacheable; every turn resends (full history + new message) as input,
// matching the fact that a stateless messages array is resent in full every turn.
function simulateSession(opts: {
  systemTokens: number; userTokensPerTurn: number; assistantTokensPerTurn: number;
  turns: number; prefixCaching: boolean; minCacheable: number; // vendor minimum cacheable prefix (e.g. 1024)
}): SessionBill {
  const { systemTokens, userTokensPerTurn, assistantTokensPerTurn, turns, prefixCaching, minCacheable } = opts;
  const bills: TurnBill[] = [];
  let historyTokens = 0;
  for (let turn = 1; turn <= turns; turn++) {
    const newTokens = userTokensPerTurn;
    const totalInput = systemTokens + historyTokens + newTokens;
    // the cache only covers the stable prefix (system prompt) and only above the minimum cacheable length
    const prefixCacheable = prefixCaching && systemTokens >= minCacheable;
    const cachedTokens = prefixCacheable ? systemTokens : 0;
    const uncachedTokens = totalInput - cachedTokens;
    bills.push({ turn, cacheHit: cachedTokens > 0, uncachedTokens, cachedTokens, cost: costOf(uncachedTokens, cachedTokens, PRICE) });
    historyTokens += newTokens + assistantTokensPerTurn;
  }
  const totalCost = bills.reduce((s, b) => s + b.cost, 0);
  const allInput = bills.reduce((s, b) => s + b.uncachedTokens + b.cachedTokens, 0);
  const allCached = bills.reduce((s, b) => s + b.cachedTokens, 0);
  return { turns: bills, totalCost, cacheHitRatio: allCached / allInput };
}

const SESSION = { systemTokens: 20_000, userTokensPerTurn: 500, assistantTokensPerTurn: 500, turns: 6, minCacheable: 1024 };

const withCache = simulateSession({ ...SESSION, prefixCaching: true });
const noCache = simulateSession({ ...SESSION, prefixCaching: false });
// negative case: system prompt below the minimum cacheable length -> the structure silently degrades
const tooShort = simulateSession({ ...SESSION, systemTokens: 800, prefixCaching: true });

const fmtCost = (c: number) => `$${c.toFixed(4)}`;
for (const [label, bill] of [['prefix caching ON', withCache], ['prefix caching OFF', noCache], ['ON but system prompt < 1024 tokens (negative)', tooShort]] as const) {
  console.log(`--- ${label} ---`);
  for (const t of bill.turns) {
    console.log(`turn ${t.turn}: input ${fmtCost(t.cost)} (uncached ${t.uncachedTokens.toLocaleString()} + cached ${t.cachedTokens.toLocaleString()})`);
  }
  console.log(`session total: ${fmtCost(bill.totalCost)}; cache hit ratio: ${(bill.cacheHitRatio * 100).toFixed(1)}% of input tokens\n`);
}
console.log(`savings from prefix caching: ${fmtCost(noCache.totalCost - withCache.totalCost)} (${(((1 - withCache.totalCost / noCache.totalCost)) * 100).toFixed(1)}%)`);

Run and expected output:

text
$ node prefix-cache-cost.mts
--- prefix caching ON ---
turn 1: input $0.0075 (uncached 500 + cached 20,000)
turn 2: input $0.0105 (uncached 1,500 + cached 20,000)
turn 3: input $0.0135 (uncached 2,500 + cached 20,000)
turn 4: input $0.0165 (uncached 3,500 + cached 20,000)
turn 5: input $0.0195 (uncached 4,500 + cached 20,000)
turn 6: input $0.0225 (uncached 5,500 + cached 20,000)
session total: $0.0900; cache hit ratio: 87.0% of input tokens

--- prefix caching OFF ---
turn 1: input $0.0615 (uncached 20,500 + cached 0)
...
turn 6: input $0.0765 (uncached 25,500 + cached 0)
session total: $0.4140; cache hit ratio: 0.0% of input tokens

--- ON but system prompt < 1024 tokens (negative) ---
turn 1: input $0.0039 (uncached 1,300 + cached 0)
...
turn 6: input $0.0189 (uncached 5,300 + cached 0)
session total: $0.0684; cache hit ratio: 0.0% of input tokens

savings from prefix caching: $0.3240 (78.3%)

(Full output is 18 lines; middle turns elided here.)

Read the negative case this way: in the third section the cache switch is ON but the hit ratio is 0% — the system prompt is below the minimum cacheable length, and the requests still succeed without an error. "Caching is enabled" is not "caching is working"; the only trustworthy evidence is the hit field in the response usage.

Acceptance: the three hit ratios read 87.0% / 0.0% / 0.0%, savings 78.3%. Cleanup: delete the file.

Scenario matrix ​

ScenarioInputActionOutputFitsDoes not fit
Basic: session cost accountingsystem prompt + turn countRun ON and OFFSession-level savings ratioCost estimation for multi-turn productsSingle one-shot calls
Common: prompt reorder payoffDynamic content mixed into the frontRestructure and re-runHit ratio changePrompt revision reviewsOutput-length problems (decode-side)
Combined: tier change estimateChange the two-price ratio or minimum lengthRe-run and compareCost under the new tierRe-checking after a vendor pricing changeNegotiating discounts (talk to sales)

3. Principles ​

The physics behind the four levers ​

Quantization. The bulk of inference cost is the bandwidth of "moving weights from memory into compute cores", not floating-point arithmetic itself. Compressing weights from FP16 to INT8/INT4 halves and halves the bytes moved — smaller model replicas, denser batches, cheaper unit prices. The trade is accuracy: whether quantization error is acceptable must be answered by your eval set, not by a vendor's "near-lossless" marketing line. Math and implementation → Learn LLM chapter 9.

Paged KV. If the KV cache is pre-allocated contiguously per request at "maximum context", fragmentation and reservation gut memory utilization. PagedAttention slices KV into fixed blocks mapped on demand (the operating-system paging idea); memory waste shrinks and one card holds more concurrent sequences — throughput rises, per-request cost falls. vLLM homepage features: "Continuous batching of incoming requests, chunked prefill, prefix caching" (retrievedAt 2026-09-01).

Prefix caching. When the same prefix appears a second time, prefill has already been computed and the KV is already in memory — skipping recomputation is "reusing an existing result". OpenAI's documentation states the mechanism plainly: requests are routed to machines that "recently processed the same prompt", and what is cached is "the key/value tensors from the model's attention layers produced during prefill" (retrievedAt 2026-09-01). The compute the vendor saves is passed back as a discount — that is the causal chain behind cheaper cached input. It lowers TTFT (prefill skipped) and cost at once.

Speculative decoding. Decode serialization is the root of the latency floor; let a small model draft several tokens and the large model verify them in one parallel pass, accepting the legal prefix — several tokens per step. The output distribution is unchanged while latency drops (vLLM supports n-gram, EAGLE, and other variants, retrievedAt 2026-09-01).

Mechanism → billing field mapping (core deliverable) ​

MechanismOpenAI sideAnthropic sideYour lever
Prefix caching (hit)The Cached input pricing column; usage prompt_tokens_details.cached_tokenscache_read_input_tokens; the Cache Hits & Refreshes pricing columnStable prefix first, dynamic last; keep turn gaps inside the TTL
Prefix caching (write)No extra fee (automatic caching, retrievedAt 2026-09-01)cache_creation_input_tokens; 5m writes at +25%, 1h writes at 2× (retrievedAt 2026-09-01)High-frequency sessions use the short TTL; consider 1h only when gaps exceed 5 minutes
Quantization / paged KV / speculative decodingNot exposed line-by-line; folded into unit prices and tiersSamePick tiers and models; configure the engine directly when self-hosting
Batch discountBatch API discount (pricing page notes it suits non-time-sensitive work)Batch API at 50% off input and outputRoute non-realtime tasks through the batch channel

Key differences between the two caching models (structure comparison, retrievedAt 2026-09-01):

DimensionOpenAIAnthropic
TriggerAutomatic (prefixes ≥1024 tokens)Explicit cache_control breakpoints (minimum 1024; 2048 for Haiku lines)
Write feeNo extra fee5m write at 1.25×, 1h write at 2× base price
Hit feeThe Cached input column price10% of base price
InvalidationPrefix mismatch; eviction after 5–10 minutes idle (up to 1 hour; extended retention up to 24 hours)Prefix mismatch (tools→system→messages hierarchy invalidates downstream); TTL expiry (a hit refreshes for free)

Spec claims vs local measurement ​

ClaimSourceLocal fixture measurement (above)
Input tokens have two prices: cached cheaper than uncachedBoth pricing pages' structure (L0)Measured: under the two-price table, ON saves 78.3% vs OFF
Hits need a matching prefix above the minimum lengthBoth caching docs (L0)Measured: systemTokens < minCacheable gives a 0% hit ratio with no request error
A stateless API resends the whole history every turnModel API ContractMeasured: turn N's uncached tokens grow linearly with turns
Anthropic writes carry a premium; hits within the 5m TTL refresh freeAnthropic caching docs (L0)Not modeled: the fixture ignores the write premium (see open questions)
A hit lowers TTFT and cost togetherCorollary of skipping prefillNot covered: the fixture bills money, not latency

Boundary with Learn LLM ​

Coding details of quantization formats (GPTQ/AWQ/GGUF etc.), methods for measuring accuracy loss, and block-management implementations of the KV cache → Learn LLM chapter 9. This page keeps only decision depth: "which lever moves which cost line".

4. Development ​

Cache-friendly engineering checklist ​

  • Freeze the prompt structure: system prompt, tool definitions, and few-shot examples at the very front; user data, timestamps, and other dynamic content at the end.
  • Instrument the hit rate: log cached_tokens / cache_read_input_tokens per request — it is the sentinel metric for both cost and TTFT.
  • Session cadence: keep consecutive turn gaps inside the TTL (Anthropic: within 5 minutes, refreshed on hit); sessions that resume much later accept misses or evaluate a longer TTL.
  • Keep the prefix byte-identical: toggles on tools, image parameters, or tool_choice changes invalidate the prefix (the invalidation table in Anthropic's docs, retrievedAt 2026-09-01).

Debug runbooks ​

Symptom → Evidence → Action → Done when ​

Symptom: the bill did not drop even though "caching is already enabled". Evidence: usage hit fields: cached_tokens / cache_read_input_tokens stuck at 0 or suddenly fallen. Action: check in order — ① is the prefix broken (dynamic content moved forward, a template version string baked into the system prompt); ② is the prefix below the minimum cacheable length; ③ do request gaps exceed the TTL; ④ do tool definitions or image parameters change every turn. Done when: the hit ratio returns to baseline and the cost curve visibly falls.

Symptom → Evidence → Action → Done when ​

Symptom: unexpected "write" charges appear on the Anthropic bill. Evidence: a high cache_creation_input_tokens share; a write/read ratio out of balance. Action: check the TTL tier (for frequent hits the 5m write is cheaper; use 1h only when gaps exceed 5 minutes); reduce unnecessary breakpoint churn; move low-frequency tasks to the Batch API. Done when: the write/read ratio matches request cadence; per-session unit cost falls.

Symptom → Evidence → Action → Done when ​

Symptom: after switching to a "quantized / small" model, it got cheaper but answer quality slipped. Evidence: fixed-eval-set scores comparing the quantized tier against the original (not gut feel). Action: roll back the precision tier; or hybrid routing — simple requests to quantized/small models, complex ones to full precision (the routing threshold goes through eval). Done when: eval scores return inside the gate; the savings survive and are documented.

Anti-patterns ​

  • Treating "the cache switch is on" as the success bar — the only evidence is the usage hit field.
  • Baking timestamps or random IDs into the front of the system prompt — every request is a miss.
  • Substituting a vendor's "near-lossless" line for your own eval-set verification of a quantization tier.
  • Watching unit prices while ignoring structure: a write premium, minimum length, or TTL can each turn a "discount" into "more expensive".
  • "Looks successful but the evidence is missing": cost fell but cannot be attributed to a field — it may be a traffic shift, not the optimization.

5. Resource Library ​

Four-level reading route ​

Resource table ​

NameLevelcanonical URLPurposeSupported claimNext
OpenAI Prompt Caching guideL0https://platform.openai.com/docs/guides/prompt-cachingAutomatic caching engineeringautomatic caching ≥1024 tokens; the cached_tokens field; prefill-produced KV tensors are what get cached; no extra write feeReorder your prompt
Anthropic Prompt Caching docsL0https://docs.anthropic.com/en/docs/build-with-claude/prompt-cachingExplicit caching engineeringcache_control; 5m/1h TTLs; writes 1.25×/2×, reads 10%; minimum 1024/2048Design breakpoint placement
OpenAI pricing pageL0https://platform.openai.com/docs/pricingBilling structureInput / Cached input / Output column structure; batch discount; reasoning tokens billed as outputBuild a cost model
Anthropic pricing pageL0https://docs.anthropic.com/en/docs/about-claude/pricingBilling structureBase Input / 5m Write / 1h Write / Cache Hit / Output column structureCompare write/read premiums
vLLM docs homepageL1https://docs.vllm.ai/en/latest/Self-hosted engine featuresfeature list: continuous batching / chunked prefill / prefix caching / INT8, INT4 quantization / speculative decodingDig into feature pages
Learn LLM ch. 9 · Inference & QuantizationEhttps://llm.zenheart.site/chapters/09-inference-cacheTheory derivationsQuantization and KV cache math live in Learn LLMHand-write a quantization experiment
This page's fixtureEprefix-cache-cost.mts (inline above)Zero-key verificationThe two-price billing structure; minimum length causing silent missesPlug in your project's real prices

retrievedAt: all web resources 2026-09-01.

Active falsification and open questions ​

  • The fixture ignores Anthropic's write premium and OpenAI's routing details: turn 1 may be billed at a "write price" rather than full price on a real vendor, so amounts differ slightly; the structural conclusions (two prices, minimum length, hit ratio) stand.
  • Unit prices are fictional round numbers: this verifies understanding of the billing structure; real amounts follow the vendor pricing page on the day (this repo does not copy unit prices).
  • "Hits lower TTFT too" comes from the skipped-prefill corollary and OpenAI's official "reduce latency by up to 80%" framing (retrievedAt 2026-09-01), not local measurement.
  • The acceptable threshold for quantization accuracy loss is task-dependent; this repo gives no universal number.

Where learn-ai stops / where to go next ​

This page owns "the mapping from mechanisms to the bill". Turn this knowledge into calling code → Model API Contract; overall cost governance and budgets → Cost and Performance; quantization and block-management math → Learn LLM chapter 9; quantization practice on-device → Browser and Edge Inference.

Built for frontend engineers · Powered by VitePress