Group: Advanced (bridge) | Previous group exit: run long-term with release gates and versioning | This page exit: know which engineering decision this topic affects, and when to go to Learn LLM
Multimodal (Bridge)
Bridge page: multimodal tokenization (how images become tokens), cross-modal alignment mechanics, and vision-encoder derivations live in Learn LLM (multimodal chapter). This repo keeps only the minimal usage an app engineer needs.
Minimal application-side usage
Using multimodality well in an application (vision as the example) takes four things:
- Message structure: images as content blocks (base64 / URL / Files API) alongside text — see the Model API contract.
- Cost awareness: images convert to tokens by size and enter the context budget — large images are a real cost item.
- Placement technique: images perform better before the instruction/question; label multiple images with
Image 1:,Image 2:. - Capability boundaries: spatial reasoning, precise counting, and person identification are known weak spots; high-stakes flows require human review (see Security).
Productization note: when the same image repeats across turns, use the Files API file_id instead of re-sending base64 every turn.
Pages in this section
| Page | Status | Content | Language side |
|---|---|---|---|
| Claude Vision capabilities (zh) | case | Claude vision integration details: image sources, limits, token math, prompting tips (figures per the official docs) | ZH only |
When to go to Learn LLM
- Want to know how images/audio are split into tokens and why size caps exist → Learn LLM, multimodal tokenization chapter
- Want to understand vision-language alignment training and where multimodal capability comes from → Learn LLM, cross-modal mechanics chapter
Vendor-side specifics (per-request image caps, file sizes, supported formats) follow the official docs in real time; this repo does not maintain snapshot figures.