Chat templates.
The model never sees your messages array. It sees one flat string of tokens, produced by a Jinja template that shipped alongside the weights, and that template — not your prompt, not your framework — decides what "system", "tool" and "assistant" actually mean to this particular model. It is the highest-leverage artefact in the stack and the only one nobody versions, diffs or tests, which is why a model that calls tools perfectly through one runtime quietly stops through another.
What it does, mechanically.
A chat model is still a next-token predictor over a single sequence. Roles, turns and tool calls do not exist at that level; they are a convention encoded as literal text and special tokens, and the chat template is the function that performs that encoding.
- Input: your
messagesarray, and — this is the part people miss — yourtoolsarray too. Tool definitions are not a separate API channel; they are rendered into the same string, in whatever shape the template author chose. - Output: one string, containing the model's special turn markers (
<|im_start|>,[INST],<|start_header_id|>— every family invents its own), which the tokeniser then maps to IDs. - The generation prompt: the template also appends the opening marker for the assistant's turn, so the model continues rather than predicting a new user message. Forget this flag and the model politely writes the user's next line for you.
Because the markers are literal text, the boundary between instruction and data is a string convention, not a type. That is the mechanical reason the instruction hierarchy is advisory and prompt injection is not a bug a vendor can patch: by the time the sequence reaches the weights, your system prompt and a retrieved web page are the same kind of thing, distinguished only by text the template wrote around them.
It is a contract, and nobody versions it.
The template ships with the weights but is not part of them. It is a text file in the model repository, and it has all the properties of a dependency and none of the discipline of one.
- It lives in two places at once. Modern repositories carry a standalone
chat_template.jinja; older ones carry achat_templatestring insidetokenizer_config.json; many carry both because the second was never removed. Which one is authoritative depends on the toolchain reading it, and the two dominant toolchains resolve that precedence in opposite directions — so a repository whose.jinjawas updated while the JSON entry went stale renders two different prompts depending on how you loaded it. - There is often more than one template per model. Many models ship a default template and a separate tool-use variant, selected only when a
toolsarray is present. Your tool-calling behaviour may be governed by a template you have never looked at. - Fine-tunes inherit it, correctly or not. A fine-tune published without its own template gets whatever the base model's was. If the fine-tuning data used a different conversation format — extremely common — every inference is now slightly off-distribution, and nothing reports it.
- Serving stacks override it. Every local runtime and several hosted ones let an operator supply a different template, and some substitute their own reimplementation by default. The model card's template and the one actually rendering your traffic are two separate facts.
Nothing in that list produces an error. A template mismatch is a quality regression, which is why it is usually blamed on quantisation, on the model, or on the prompt.
The four ways it goes wrong.
Ordered by how often they are misdiagnosed rather than by severity.
- Right model, wrong template. The inheritance and precedence cases above. Symptom: the model is subtly worse than its published benchmark scores, across the board, with no single reproducible failure. This is the hardest to spot because there is nothing to spot.
- Right template, different renderer. Serving stacks reimplement Jinja, or reimplement the template itself in another language for speed. Any reimplementation is a chance to diverge, and the divergences are invisible — a rewritten template that gets one separator wrong produces a fluent, plausible, slightly degraded model.
- Whitespace, which is not cosmetic. A space added by a template is not free: depending on where it lands it can be absorbed into a neighbouring token or become a token of its own, changing the sequence the model was trained to expect. This is the clearest illustration that templates operate below the level of text and belong to tokenisation, not to prompting.
- Tools rendered by a template that does not support them. If the template has no branch for the
toolsarray, your carefully designed schemas are either dropped entirely or flattened into a system message by the serving layer's fallback. The model then calls tools in whatever format it feels like, the runtime fails to parse it, and the call arrives as prose in the content field — which your framework, watching for a structuredtool_callsarray, reads as a refusal. See tool calling for the round trip this breaks, and the wiki's comparison of four local runtimes for how differently four projects handle exactly this.
All four failures share a signature: fluent output, no exception, degraded results. That makes the template the first thing to check when a model underperforms in your stack and the last thing anyone does check — because the mental model is "prompt goes in, answer comes out", and the template is invisible in that picture.
Treat it as a versioned dependency.
Four habits, in order of how much they save you.
- Print the fully rendered prompt. Every serving stack will show it to you, and it is the single highest-value debugging artefact in local inference. Do it once per model on adoption, and again whenever behaviour changes. More people have found their bug in that string than in any evaluation.
- Pin the template with the weights, and diff it on upgrade. It is a file; check it into your repository next to the model reference. A template change is a model change — it invalidates your evaluation baselines exactly as a new checkpoint would, and it silently invalidates your prompt cache, because the cached prefix no longer matches a single byte of the new one.
- Prefer constraint over parsing for structured output. A runtime that constrains generation to a grammar cannot emit a malformed tool call; a runtime that parses afterwards can only discover one. That is the practical argument for constrained decoding and the reason it matters far more when you self-host than when a provider owns both ends of the translation.
- Record which template produced a run. A hash of the rendered prefix, stored with the trace, turns "the model got worse last Tuesday" from an argument into a lookup — and it is the same discipline reproducibility asks of every other input.
If you do one thing: take the model you are running locally and dump the rendered prompt for a request that includes a tool definition. Look for three things — that your tools appear in it at all, that the turn markers match the ones on the model card, and that there is no stray whitespace before the generation prompt. That check takes five minutes, it is the only way to see the layer you are actually programming against, and it resolves a surprising share of "this model is bad at tool calling".
Related: system vs user prompts for the roles the template encodes, small & local models for where this bites hardest, and model deprecation & migration for handling the template change that arrives with the next checkpoint.