AI Blog

Ollama vs LM Studio vs llama.cpp vs MLX: the tool call is the whole difference

Four local runtimes, the same weights, four different prompts going in and four different answers to whether a tool call comes back parsed. None of that is throughput, and throughput is the only axis anyone compares.

By Agentic AI Wiki 18 min read

Point four local runtimes at the same weights, the same prompt and the same OpenAI-compatible endpoint, and you will get four different prompts on the way in and, on the way out, anywhere between a parsed tool_calls array and a blob of XML sitting in content that your agent framework will treat as an answer. Nothing in that sentence is about tokens per second, and tokens per second is the only axis anyone compares. The layer that decides whether your local agent works at all is the chat template going in and the tool-call parser coming out — and all four runtimes reimplement both, differently, from scratch.

At a glance

Two of these are engines and two are experience layers wrapped around engines, which is the usual framing and is the wrong one for an agent. The axis that matters is who builds the prompt and who parses the reply.

RuntimePrompt built byTool call parsed byConstrained output?
llama.cpp Its own C++ Jinja engine, written in-house A differential autoparser, plus 15 hand-written parsers Yes — GBNF derived from the parser
Ollama Hand-written Go renderers, or a Go text template, or llama.cpp 19 hand-written Go parsers, or a substring scan No
LM Studio Closed; or its own injected system prompt Closed; falls back to a bespoke text protocol On MLX only, via llguidance
mlx-lm The model author's own template, via Hugging Face 12 hand-written parsers, or nothing at all No
Hand-written, model-specific tool-call parsers shipped per runtime Horizontal bar chart. Ollama ships 19 parser files, llama.cpp 15, mlx-lm 12. LM Studio's parsers are closed source and are not counted. Counted from the repositories on 21 September 2026. Model-specific tool-call parsers maintained in-tree 0 5 10 15 20 ollama 19 llama.cpp 15 mlx-lm 12 LM Studio maintains its own set privately; the app is not open source, so it cannot be counted here.
Three projects independently maintaining a per-model parser zoo for the same models. LM Studio maintains a fourth, privately.
Where each local runtime is strong and weak for tool-calling agents Feature matrix with four runtimes as rows and four axes as columns. llama.cpp is medium on prompt fidelity and strong on parsing, constrained output and tool choice. Ollama is weak on prompt fidelity, medium on parsing, and weak on constrained output and tool choice. LM Studio is weak on prompt fidelity, medium on parsing and constrained output, and medium on tool choice. mlx-lm is strong on prompt fidelity and weak on parsing, constrained output and tool choice. Four axes that decide whether a local agent calls tools Prompt fidelity Tool parsing Constrained output tool_choice llama.cpp Medium Strong Strong (GBNF) Strong ollama Weak (Go rewrite) Medium None Unsupported lm studio Weak (rewrites) Medium MLX only Medium mlx-lm Strong (reference) Weak None Undocumented Strong Medium Weak or absent
No row is strong across all four. The one that is faithful on the way in is the weakest on the way out.

The layer nobody benchmarks

The two translation steps around identical weights A pipeline from a messages and tools request, through rendering the chat template, through the model weights, through parsing the tool markup, to a structured tool calls array. The render step and the parse step are highlighted as the two places each runtime substitutes its own implementation; the weights in the middle are identical across all four runtimes. messages + tools Render template Weights identical Parse markup tool_calls Substituted at the render step llama.cppits own C++ Jinja engine, plus message rewrites before rendering ollamahand-written Go renderers, or a Go text template, or delegate lm studioclosed, or an injected system prompt with a bespoke format mlx-lmthe model author's template, run by reference Jinja2 Substituted at the parse step llama.cppautoparser derived from the template, plus a grammar that constrains it ollamaper-model Go parsers, else a substring scan for the tool's name lm studioclosed; unparsed calls are returned as content mlx-lmper-model parsers, else a warning and raw text in content Both substituted steps fail silently: a bad render degrades quality, a bad parse hides the tool call in a field nothing reads.
The weights are the same in all four columns. Everything that differs sits on either side of them.

An OpenAI-shaped tool call is a round trip through two translations. Your messages and tools arrays are rendered into one flat string using a Jinja template that shipped with the model; the model emits text in whatever markup it was trained on — <tool_call>, [TOOL_CALLS], <|tool_call>, a channel header, a Python-ish call expression; and the runtime parses that text back into structured tool_calls. The weights only see the middle.

Both translations are per-model, both are undocumented in any normative sense, and both fail quietly. When the render is wrong the model receives a prompt that does not match its training and produces slightly worse tool calls — no error. When the parse is wrong the tool call arrives as prose in the content field, and your agent framework, which is looking at tool_calls, sees a model that declined to use a tool. That failure mode is why local agents "just don't follow instructions" and why the fix is almost never a bigger model. It is also why constrained decoding matters far more here than it does against a hosted API, where the provider owns both ends of the translation.

llama.cpp — from a parser zoo to a parser that reads the template

llama.cpp (MIT, 129,000 stars, semver v0.4.1 on 14 September 2026 alongside rolling build tags) has spent 2026 going in the opposite direction from everyone else, and the numbers make it vivid. Before 6 March its common_chat_format enum carried 29 values — LLAMA_3_X, HERMES_2_PRO, FUNCTIONARY_V3_2, COMMAND_R7B, DEEPSEEK_R1, one per model family, each with its own parser. Today the enum has five.

What replaced the other 24 is genuinely clever: a differential autoparser that reads the model's own chat template to work out the markup. It renders the template several times with sentinel values — the source literally contains FFF_FIRST_FUN_F and AA_ARG_FST_AA as probe strings — and diffs the outputs to extract the delimiters, the argument encoding, the call-ID position and whether reasoning is tag-based. From that it builds a PEG parser, and from the PEG parser it derives a GBNF grammar that constrains generation, triggered lazily when the model emits the tool-section marker.

That last step is the most valuable thing in this entire comparison and the only place it exists: llama.cpp is alone in constraining the shape of the tool call at sampling time rather than hoping and then parsing. A malformed tool call is not a parse failure to recover from; it is a token sequence the sampler would not emit.

The architecture does not eliminate special cases, it relocates them. Fifteen hand-written parsers still live in common/parsers/, split out in September, for the models whose markup defeats the general mechanism — including one whose namespace token collides with the autoparser's own delimiters. Detection of those cases is a chain of substring tests against the raw template text, and one of them prints a warning that tells you exactly how load-bearing the template file is: it detects an outdated Gemma 4 template and applies compatibility workarounds.

Two caveats worth knowing before you trust the fidelity. llama.cpp replaced its vendored Jinja library with an engine written from scratch in January, and that engine's own README documents two known departures from reference behaviour — whitespace added by a template is tokenised as a standalone token, and special tokens assembled dynamically from user input do not work. There is also a workaround namespace that rewrites your message array before rendering: remapping the developer role to system, backfilling non-null content, coercing tool arguments from strings to objects. All defensible; none of it byte-identical to what the model author intended. And the documentation lags the code badly — the server README still tells you function calling needs the --jinja flag, which became the default in December 2025.

Ollama — three chat paths and a substring scan

Ollama (MIT, 181,000 stars, v0.34.2 on 17 September 2026) is the most used of the four and has the most surprising internals. It vendors llama.cpp — pinned, on the day of writing, about 98 builds behind upstream — and then, for a large and growing set of models, switches llama.cpp's chat stack off entirely and does the work itself in Go.

There are three paths. Models with an Ollama-specific renderer or parser in their manifest get Go code on both ends, with DisableJinja set. Older models get a Go text/template plus a generic parser. Everything else is handed to the embedded llama-server, which does what the previous section describes. Which path a model takes is a field in its manifest — and in at least one family, a filename heuristic that looks for e2b or 12b in the model name to pick between a small and a large renderer variant.

The reimplementation is substantial: around 20 renderer files covering 26 model names, and 19 parser files covering 27. Every one of those is a second implementation of something the model already shipped as a Jinja template, maintained by a different team, reached by a different code path. Ollama does cross-check about eight of them byte-for-byte against the model author's real template — the reference templates are checked into the repo — but that test is opt-in behind an environment variable and covers under a third of the renderers.

The generic fallback is where it gets uncomfortable. To find a model's tool-call marker, Ollama walks the Go template's AST, locates the {{if .ToolCalls}} node, takes the first text that follows it, and uses that as the opening tag — and if it finds nothing, the code comment says it returns "{" to indicate that JSON objects should be attempted as tool calls. Detection of a specific tool is then a byte-level substring search for the tool's own name in the model's output. It works more often than it sounds like it should, and its failures are exactly the ones you would predict: the open issue list includes tool calls silently discarded when parsing fails, content after a tool call being dropped, and a tool call vanishing entirely when a parameter happens to be named name.

There is no grammar anywhere in Ollama. A repository-wide search for grammar or GBNF returns a single comment. It renders, generates freely, and parses what comes back. It also documents tool_choice as unsupported on both its OpenAI- and Anthropic-compatible endpoints, which rules it out as a backend for any agent that needs to force a call.

LM Studio — two tiers, and a protocol nobody else speaks

LM Studio is the only closed-source subject here; its SDKs and CLI are MIT but the application is not published. It runs both llama.cpp and, on Apple silicon, MLX, swapping engines as hot-loadable runtimes.

Its tool-use model is explicitly two-tier and the documentation is unusually honest about it. "Native" support requires two things at once: the model's chat template supports tools, and LM Studio has implemented that model's specific format. Models meeting both get a hammer badge in the UI. The published list of natively supported families is three — Qwen, Llama 3.1/3.2, and Mistral — and its accompanying log excerpt is dated 2024, so treat the list as indicative rather than current.

Everything else falls to "Default" tool use, and this is the part with no analogue in the other three. LM Studio injects its own system prompt instructing the model to emit calls in a format it invented:

[TOOL_REQUEST]{"name": "tool_name", "arguments": {"param1": "value1"}}[END_TOOL_REQUEST]

It also rewrites the conversation to fit templates that lack a tool role — converting tool-role messages to user, and rewriting prior assistant tool_calls into the same bespoke syntax. The documentation notes the format is subject to change, and that if it cannot parse a correctly formatted call it returns the text in content.

Read that as an engineering position rather than a flaw: LM Studio has decided that a uniform format the model has never seen, taught in-context, beats a per-model format it has not implemented. For a chat application that is a reasonable trade. For an agent it means your tool-calling reliability depends on in-context instruction-following by a model that was fine-tuned on completely different markup, which is a much weaker guarantee than either of the alternatives — and the bug tracker carries a steady stream of exactly that symptom, tool calls arriving as raw text with tool_calls empty.

The exception is worth noting because it is the second-best mechanism in this article: LM Studio's MLX engine is open source, and it constrains tool output with llguidance against hardcoded per-model markers. On Apple silicon with an MLX model, LM Studio does the thing Ollama and mlx-lm do not.

mlx-lm — a faithful prompt and, sometimes, no parser at all

Apple's mlx-lm is the smallest project here (7,000 stars against llama.cpp's 129,000) and it wins the half of the problem that everyone else loses. It calls Hugging Face transformers' apply_chat_template directly, which means the prompt your model sees is the model author's template, executed by reference Jinja2, with no reimplementation in between. Nothing else in this comparison can say that.

The reply side is the inverse. There are 12 tool parsers, selected by substring-matching the chat template — "<minimax:tool_call>" in chat_template, "[TOOL_CALLS]" in chat_template, and so on — with a fallback to inspecting the vocabulary. The generic parser, for models whose markup is a plain <tool_call> wrapper, is three lines: strip, json.loads, return.

And when nothing matches, mlx-lm does not refuse. It logs a warning that the model does not support tool calling, then proceeds — passing your tools array into the template anyway, so the model is told about the tools, calls one, and the call comes back as text in content with no tool_calls field and finish_reason unchanged. That is the cleanest demonstration available of this article's thesis: identical weights, a more faithful prompt than any competitor produced, and the tool call lost entirely at the parse layer. Two further practicalities: the server's own documentation does not mention the tools parameter at all despite implementing it, and the last PyPI release predates the current main branch by about five months, so anyone pinning from PyPI is running noticeably older parsing than the repository suggests.

The bug that proves the point: two projects read the same files in opposite order

Two toolchains, opposite precedence for the same two template files One model repository containing both a chat_template.jinja file and a chat_template entry inside tokenizer_config.json. The Hugging Face transformers path prefers the standalone jinja file; the GGUF conversion path used by llama.cpp, Ollama and LM Studio prefers the tokenizer_config entry and treats the jinja file as a fallback. The result is two different rendered prompts from one repository, with no error or warning. One repository ships both files. Two toolchains disagree about which one wins. chat_template.jinja tokenizer_config.json transformers / mlx-lm GGUF conversion — llama.cpp, Ollama, LM Studio Standalone .jinja file takes priority tokenizer_config entry takes priority The current template Whatever the legacy string still says Same model, two prompts, no error — you will read it as a quantisation problem
Same repository, same model, two rendered prompts — decided by a precedence rule that the two toolchains define in opposite directions.

A model repository can carry its chat template in two places: a standalone chat_template.jinja file, and a chat_template string inside tokenizer_config.json. Many carry both, because the file is the newer convention and the JSON entry was left behind.

Hugging Face transformers resolves this explicitly, with a comment in the source saying that independent chat-template files take priority over template entries in the tokenizer config. llama.cpp's GGUF converter resolves it the other way: it reads tokenizer_config.json's chat_template first and uses the .jinja file only as a fallback.

So for any model whose .jinja file has been updated while the JSON entry went stale, the prompt rendered by mlx-lm and the prompt baked into every GGUF — and therefore served by llama.cpp, Ollama and LM Studio — come from different templates. No error, no warning, no version mismatch to notice. You will observe it as a model that is mysteriously worse at tool calling in GGUF form than the benchmark numbers suggested, and you will blame the quantisation.

This is what "the template is an unversioned string contract" means in practice. Nobody signs it, nobody diffs it, two major toolchains disagree about which copy is authoritative, and the failure is a quality regression rather than an exception. If you are debugging a local agent that calls tools badly, dump the fully rendered prompt before you change anything else — every one of these runtimes will show it to you, and it is the fastest way to find out you have been testing a different prompt than you thought.

When to pick which

If you are…PickBecause
Building an agent that must call tools reliablyllama.cpp directlyIt is the only one that constrains the tool call at sampling time instead of parsing afterwards, and it exposes the template and its parsing decisions for inspection.
On Apple silicon and want both halvesLM Studio with the MLX engineReference-quality MLX prompting plus llguidance-constrained output — the only combination that gets both ends, at the cost of a closed app.
Running evaluations or comparing against published numbersmlx-lmIt is the only one that renders the model author's template unmodified, which is what the published scores were produced with. Check it has a parser for your model first.
Standing up many models quickly for non-agentic workOllamaThe model library and pull-and-run ergonomics are genuinely unmatched, and none of this matters when you are not parsing tool calls.
Deciding on tokens per secondAny of themThey share engines. Whichever one you pick, the thing that will break your agent is on this page and not in a throughput chart.

Before you choose anything, run the same tool-calling request against two runtimes serving the same weights and diff three fields: the rendered prompt, the raw completion text, and whether tool_calls came back populated. Twenty minutes. If the prompts differ you have found your quality gap; if the completions are similar but only one produced tool_calls, you have found a parser gap — and those are the only two bugs in this entire category.

FAQ

Do these runtimes produce different output from the same model?

Yes, and in two distinct ways. They build different prompts from the same messages array, because three of the four reimplement the chat template rather than running the model author's. And they parse the model's output differently, so a tool call that one returns as a structured tool_calls array another returns as plain text in content.

Which one is most faithful to the model author's intent?

mlx-lm on the prompt side, because it calls Hugging Face's apply_chat_template directly. No runtime here guarantees byte-identical rendering, and mlx-lm has the weakest parsing of the four, so faithfulness on the way in does not by itself make it the best agent backend.

Why does my local model "refuse to use tools" when the hosted version does?

The most likely explanation is that it did use a tool and your runtime failed to parse the call, leaving it as text in content where your framework never looks. Check the raw completion before concluding anything about the model.

Does any of them constrain tool-call output with a grammar?

llama.cpp does, deriving a GBNF grammar from its parser and triggering it lazily on the tool-section marker. LM Studio's MLX engine does, using llguidance. Ollama has no grammar code at all, and mlx-lm has none for tools.

Is Ollama just a wrapper around llama.cpp?

Less than it used to be. It vendors llama.cpp and MLX, but for a growing list of models it disables llama.cpp's chat handling and renders the prompt and parses the reply in its own Go code — around 26 model names covered by its own renderers and 27 by its own parsers at the time of writing.

How much does tokens per second actually matter here?

Less than the comparison articles imply, because these projects share engines — Ollama and LM Studio both run llama.cpp, and both increasingly run MLX on Apple silicon. Throughput differences are mostly configuration. Tool-calling fidelity differences are architectural.

Further reading

On this wiki:

Project sources: