Point four local runtimes at the same weights, the same prompt and the same OpenAI-compatible endpoint, and you will get four different prompts on the way in and, on the way out, anywhere between a parsed tool_calls array and a blob of XML sitting in content that your agent framework will treat as an answer. Nothing in that sentence is about tokens per second, and tokens per second is the only axis anyone compares. The layer that decides whether your local agent works at all is the chat template going in and the tool-call parser coming out — and all four runtimes reimplement both, differently, from scratch.
At a glance
Two of these are engines and two are experience layers wrapped around engines, which is the usual framing and is the wrong one for an agent. The axis that matters is who builds the prompt and who parses the reply.
| Runtime | Prompt built by | Tool call parsed by | Constrained output? |
|---|---|---|---|
| llama.cpp | Its own C++ Jinja engine, written in-house | A differential autoparser, plus 15 hand-written parsers | Yes — GBNF derived from the parser |
| Ollama | Hand-written Go renderers, or a Go text template, or llama.cpp | 19 hand-written Go parsers, or a substring scan | No |
| LM Studio | Closed; or its own injected system prompt | Closed; falls back to a bespoke text protocol | On MLX only, via llguidance |
| mlx-lm | The model author's own template, via Hugging Face | 12 hand-written parsers, or nothing at all | No |
The layer nobody benchmarks
An OpenAI-shaped tool call is a round trip through two translations. Your messages and tools arrays are rendered into one flat string using a Jinja template that shipped with the model; the model emits text in whatever markup it was trained on — <tool_call>, [TOOL_CALLS], <|tool_call>, a channel header, a Python-ish call expression; and the runtime parses that text back into structured tool_calls. The weights only see the middle.
Both translations are per-model, both are undocumented in any normative sense, and both fail quietly. When the render is wrong the model receives a prompt that does not match its training and produces slightly worse tool calls — no error. When the parse is wrong the tool call arrives as prose in the content field, and your agent framework, which is looking at tool_calls, sees a model that declined to use a tool. That failure mode is why local agents "just don't follow instructions" and why the fix is almost never a bigger model. It is also why constrained decoding matters far more here than it does against a hosted API, where the provider owns both ends of the translation.
llama.cpp — from a parser zoo to a parser that reads the template
llama.cpp (MIT, 129,000 stars, semver v0.4.1 on 14 September 2026 alongside rolling build tags) has spent 2026 going in the opposite direction from everyone else, and the numbers make it vivid. Before 6 March its common_chat_format enum carried 29 values — LLAMA_3_X, HERMES_2_PRO, FUNCTIONARY_V3_2, COMMAND_R7B, DEEPSEEK_R1, one per model family, each with its own parser. Today the enum has five.
What replaced the other 24 is genuinely clever: a differential autoparser that reads the model's own chat template to work out the markup. It renders the template several times with sentinel values — the source literally contains FFF_FIRST_FUN_F and AA_ARG_FST_AA as probe strings — and diffs the outputs to extract the delimiters, the argument encoding, the call-ID position and whether reasoning is tag-based. From that it builds a PEG parser, and from the PEG parser it derives a GBNF grammar that constrains generation, triggered lazily when the model emits the tool-section marker.
That last step is the most valuable thing in this entire comparison and the only place it exists: llama.cpp is alone in constraining the shape of the tool call at sampling time rather than hoping and then parsing. A malformed tool call is not a parse failure to recover from; it is a token sequence the sampler would not emit.
The architecture does not eliminate special cases, it relocates them. Fifteen hand-written parsers still live in common/parsers/, split out in September, for the models whose markup defeats the general mechanism — including one whose namespace token collides with the autoparser's own delimiters. Detection of those cases is a chain of substring tests against the raw template text, and one of them prints a warning that tells you exactly how load-bearing the template file is: it detects an outdated Gemma 4 template and applies compatibility workarounds.
Two caveats worth knowing before you trust the fidelity. llama.cpp replaced its vendored Jinja library with an engine written from scratch in January, and that engine's own README documents two known departures from reference behaviour — whitespace added by a template is tokenised as a standalone token, and special tokens assembled dynamically from user input do not work. There is also a workaround namespace that rewrites your message array before rendering: remapping the developer role to system, backfilling non-null content, coercing tool arguments from strings to objects. All defensible; none of it byte-identical to what the model author intended. And the documentation lags the code badly — the server README still tells you function calling needs the --jinja flag, which became the default in December 2025.
Ollama — three chat paths and a substring scan
Ollama (MIT, 181,000 stars, v0.34.2 on 17 September 2026) is the most used of the four and has the most surprising internals. It vendors llama.cpp — pinned, on the day of writing, about 98 builds behind upstream — and then, for a large and growing set of models, switches llama.cpp's chat stack off entirely and does the work itself in Go.
There are three paths. Models with an Ollama-specific renderer or parser in their manifest get Go code on both ends, with DisableJinja set. Older models get a Go text/template plus a generic parser. Everything else is handed to the embedded llama-server, which does what the previous section describes. Which path a model takes is a field in its manifest — and in at least one family, a filename heuristic that looks for e2b or 12b in the model name to pick between a small and a large renderer variant.
The reimplementation is substantial: around 20 renderer files covering 26 model names, and 19 parser files covering 27. Every one of those is a second implementation of something the model already shipped as a Jinja template, maintained by a different team, reached by a different code path. Ollama does cross-check about eight of them byte-for-byte against the model author's real template — the reference templates are checked into the repo — but that test is opt-in behind an environment variable and covers under a third of the renderers.
The generic fallback is where it gets uncomfortable. To find a model's tool-call marker, Ollama walks the Go template's AST, locates the {{if .ToolCalls}} node, takes the first text that follows it, and uses that as the opening tag — and if it finds nothing, the code comment says it returns "{" to indicate that JSON objects should be attempted as tool calls. Detection of a specific tool is then a byte-level substring search for the tool's own name in the model's output. It works more often than it sounds like it should, and its failures are exactly the ones you would predict: the open issue list includes tool calls silently discarded when parsing fails, content after a tool call being dropped, and a tool call vanishing entirely when a parameter happens to be named name.
There is no grammar anywhere in Ollama. A repository-wide search for grammar or GBNF returns a single comment. It renders, generates freely, and parses what comes back. It also documents tool_choice as unsupported on both its OpenAI- and Anthropic-compatible endpoints, which rules it out as a backend for any agent that needs to force a call.
LM Studio — two tiers, and a protocol nobody else speaks
LM Studio is the only closed-source subject here; its SDKs and CLI are MIT but the application is not published. It runs both llama.cpp and, on Apple silicon, MLX, swapping engines as hot-loadable runtimes.
Its tool-use model is explicitly two-tier and the documentation is unusually honest about it. "Native" support requires two things at once: the model's chat template supports tools, and LM Studio has implemented that model's specific format. Models meeting both get a hammer badge in the UI. The published list of natively supported families is three — Qwen, Llama 3.1/3.2, and Mistral — and its accompanying log excerpt is dated 2024, so treat the list as indicative rather than current.
Everything else falls to "Default" tool use, and this is the part with no analogue in the other three. LM Studio injects its own system prompt instructing the model to emit calls in a format it invented:
[TOOL_REQUEST]{"name": "tool_name", "arguments": {"param1": "value1"}}[END_TOOL_REQUEST]
It also rewrites the conversation to fit templates that lack a tool role — converting tool-role messages to user, and rewriting prior assistant tool_calls into the same bespoke syntax. The documentation notes the format is subject to change, and that if it cannot parse a correctly formatted call it returns the text in content.
Read that as an engineering position rather than a flaw: LM Studio has decided that a uniform format the model has never seen, taught in-context, beats a per-model format it has not implemented. For a chat application that is a reasonable trade. For an agent it means your tool-calling reliability depends on in-context instruction-following by a model that was fine-tuned on completely different markup, which is a much weaker guarantee than either of the alternatives — and the bug tracker carries a steady stream of exactly that symptom, tool calls arriving as raw text with tool_calls empty.
The exception is worth noting because it is the second-best mechanism in this article: LM Studio's MLX engine is open source, and it constrains tool output with llguidance against hardcoded per-model markers. On Apple silicon with an MLX model, LM Studio does the thing Ollama and mlx-lm do not.
mlx-lm — a faithful prompt and, sometimes, no parser at all
Apple's mlx-lm is the smallest project here (7,000 stars against llama.cpp's 129,000) and it wins the half of the problem that everyone else loses. It calls Hugging Face transformers' apply_chat_template directly, which means the prompt your model sees is the model author's template, executed by reference Jinja2, with no reimplementation in between. Nothing else in this comparison can say that.
The reply side is the inverse. There are 12 tool parsers, selected by substring-matching the chat template — "<minimax:tool_call>" in chat_template, "[TOOL_CALLS]" in chat_template, and so on — with a fallback to inspecting the vocabulary. The generic parser, for models whose markup is a plain <tool_call> wrapper, is three lines: strip, json.loads, return.
And when nothing matches, mlx-lm does not refuse. It logs a warning that the model does not support tool calling, then proceeds — passing your tools array into the template anyway, so the model is told about the tools, calls one, and the call comes back as text in content with no tool_calls field and finish_reason unchanged. That is the cleanest demonstration available of this article's thesis: identical weights, a more faithful prompt than any competitor produced, and the tool call lost entirely at the parse layer. Two further practicalities: the server's own documentation does not mention the tools parameter at all despite implementing it, and the last PyPI release predates the current main branch by about five months, so anyone pinning from PyPI is running noticeably older parsing than the repository suggests.
The bug that proves the point: two projects read the same files in opposite order
A model repository can carry its chat template in two places: a standalone chat_template.jinja file, and a chat_template string inside tokenizer_config.json. Many carry both, because the file is the newer convention and the JSON entry was left behind.
Hugging Face transformers resolves this explicitly, with a comment in the source saying that independent chat-template files take priority over template entries in the tokenizer config. llama.cpp's GGUF converter resolves it the other way: it reads tokenizer_config.json's chat_template first and uses the .jinja file only as a fallback.
So for any model whose .jinja file has been updated while the JSON entry went stale, the prompt rendered by mlx-lm and the prompt baked into every GGUF — and therefore served by llama.cpp, Ollama and LM Studio — come from different templates. No error, no warning, no version mismatch to notice. You will observe it as a model that is mysteriously worse at tool calling in GGUF form than the benchmark numbers suggested, and you will blame the quantisation.
This is what "the template is an unversioned string contract" means in practice. Nobody signs it, nobody diffs it, two major toolchains disagree about which copy is authoritative, and the failure is a quality regression rather than an exception. If you are debugging a local agent that calls tools badly, dump the fully rendered prompt before you change anything else — every one of these runtimes will show it to you, and it is the fastest way to find out you have been testing a different prompt than you thought.
When to pick which
| If you are… | Pick | Because |
|---|---|---|
| Building an agent that must call tools reliably | llama.cpp directly | It is the only one that constrains the tool call at sampling time instead of parsing afterwards, and it exposes the template and its parsing decisions for inspection. |
| On Apple silicon and want both halves | LM Studio with the MLX engine | Reference-quality MLX prompting plus llguidance-constrained output — the only combination that gets both ends, at the cost of a closed app. |
| Running evaluations or comparing against published numbers | mlx-lm | It is the only one that renders the model author's template unmodified, which is what the published scores were produced with. Check it has a parser for your model first. |
| Standing up many models quickly for non-agentic work | Ollama | The model library and pull-and-run ergonomics are genuinely unmatched, and none of this matters when you are not parsing tool calls. |
| Deciding on tokens per second | Any of them | They share engines. Whichever one you pick, the thing that will break your agent is on this page and not in a throughput chart. |
Before you choose anything, run the same tool-calling request against two runtimes serving the same weights and diff three fields: the rendered prompt, the raw completion text, and whether tool_calls came back populated. Twenty minutes. If the prompts differ you have found your quality gap; if the completions are similar but only one produced tool_calls, you have found a parser gap — and those are the only two bugs in this entire category.
FAQ
Do these runtimes produce different output from the same model?
Yes, and in two distinct ways. They build different prompts from the same messages array, because three of the four reimplement the chat template rather than running the model author's. And they parse the model's output differently, so a tool call that one returns as a structured tool_calls array another returns as plain text in content.
Which one is most faithful to the model author's intent?
mlx-lm on the prompt side, because it calls Hugging Face's apply_chat_template directly. No runtime here guarantees byte-identical rendering, and mlx-lm has the weakest parsing of the four, so faithfulness on the way in does not by itself make it the best agent backend.
Why does my local model "refuse to use tools" when the hosted version does?
The most likely explanation is that it did use a tool and your runtime failed to parse the call, leaving it as text in content where your framework never looks. Check the raw completion before concluding anything about the model.
Does any of them constrain tool-call output with a grammar?
llama.cpp does, deriving a GBNF grammar from its parser and triggering it lazily on the tool-section marker. LM Studio's MLX engine does, using llguidance. Ollama has no grammar code at all, and mlx-lm has none for tools.
Is Ollama just a wrapper around llama.cpp?
Less than it used to be. It vendors llama.cpp and MLX, but for a growing list of models it disables llama.cpp's chat handling and renders the prompt and parses the reply in its own Go code — around 26 model names covered by its own renderers and 27 by its own parsers at the time of writing.
How much does tokens per second actually matter here?
Less than the comparison articles imply, because these projects share engines — Ollama and LM Studio both run llama.cpp, and both increasingly run MLX on Apple silicon. Throughput differences are mostly configuration. Tool-calling fidelity differences are architectural.
Further reading
On this wiki:
- Chat templates — the unversioned string contract underneath all of this.
- Small & local models — when running locally is the right call in the first place.
- Constrained decoding — why grammar-constrained tool calls beat parsing afterwards.
- On-device agent architecture — the wider shape of an agent that runs on the user's machine.
- vLLM vs SGLang vs TensorRT-LLM vs llama.cpp — the same engines judged on serving throughput instead.