Search

Fix Broken Tool Calls in vLLM, SGLang and llama.cpp

Malformed JSON, stalled generation, hard crashes. Three causes across four engines, an ordered bisect costing one restart each, and the 21 issues behind them.

Mohit5 min read

Filed under Local LLM Inference· see every report on this topic

Three causes of tool-call corruption converging on one malformed structured call.

Verdict: Log the raw completion beside the parsed tool call, then restart without speculative decoding. That costs one restart and about twenty minutes. It identifies or rules out the two common causes before you waste time swapping models or engines.

The operator and the failure signal

This is for the operator of a self-hosted agent lane where tools stop being called, arrive as malformed JSON or empty strings, stall before the call, or collapse into repeated tokens and filler reasoning. A healthy tokens-per-second benchmark does not see any of this; it cannot tell whether the tokens formed a tool call.

Separate parser failures from decoding failures

Issue bodies, not titles, were re-verified on 2026-08-22. Several citations from an earlier draft were removed because their bodies concerned a VRAM leak, throughput complaint, or config question instead of tool-call corruption. The remaining reports are user reports and open issues, not confirmed vendor behaviour.

Cause Lives in Reports
Parser bug Engine extraction of a call from raw output vLLM #53246, #53227, #45167, #44676; SGLang #35565, #35564; llama.cpp #26987, #27363; Ollama #17638, #16648, #16686
Speculative decoding Draft-and-verify decode path vLLM #46249, #36872, #34449, #43221; SGLang #32038, #9187
Quantization Weight precision vLLM #13530
Both stacked Interaction, not sum vLLM #36872’s AWQ-4bit note; SGLang #4351, #35324

Eleven reports across four engines that share no parser code point to the first row. The parser, not the model, turns raw output into a structured call. Quantization and speculative decoding can also compound: reports cover malformed JSON, stalled generation, and hard crashes when the two are stacked.

Speculative decoding proposes future tokens through a cheap draft path, verifies them against the target model, and keeps matches. Quantization shifts those distributions slightly. That is usually an acceptable trade, but small shifts matter when the other mechanism makes accept-or-reject decisions from the same distributions.

Run this one-change-per-restart bisect

Use the same prompt and realistic concurrency at every step. Do not combine changes.

1. Capture raw output and the parsed call

This distinguishes parser output from upstream decoding. Prerequisite: the client sends tools and VLLM_URL points at the server.

# vLLM: capture the raw completion before tool-call parsing.
# Prerequisite: your client sends tools; VLLM_URL points at the server.
curl -s "$VLLM_URL/v1/chat/completions" \
  -H 'Content-Type: application/json' \
  -d '{"model":"'"$MODEL"'","messages":[{"role":"user","content":"What is the weather in Paris?"}],
       "tools":[{"type":"function","function":{"name":"get_weather",
                 "parameters":{"type":"object","properties":{"city":{"type":"string"}}}}}],
       "tool_choice":"auto","max_tokens":128}' \
| python3 -m json.tool

Read choices[0].message. Well-formed call text in content with an empty or wrong tool_calls result means a parser problem. Malformed content is upstream in decoding or the model.

2. Disable speculative decoding

Restart without draft/MTP configuration. In vLLM, remove --speculative-config, or the older --speculative-model / --num-speculative-tokens flags depending on version. In SGLang, remove the --speculative-algorithm family. If calls recover, you have a throughput trade to make.

3. Try unquantized weights

Do this only if VRAM permits it. If it fixes the problem, you are in vLLM #13530 territory: choose a quantization format rather than deciding whether quantization is allowed.

4. Change the parser before the engine

Most engines offer a parser choice. In vLLM, test --tool-call-parser with its matching --chat-template. This is cheaper than an engine migration and resolves a surprising share of cases.

5. Change engine or model last

At this point you have eliminated three cheaper causes. Starting here turns a bad afternoon into a bad week.

Keep the reproduction honest

On vLLM 0.24.0 with Qwen3.8-27B BF16 and MTP=3, concurrent long-context tool work produced repeated ! reasoning, empty assistant messages, and finish_reason=length before one tool call. Restarting without MTP fixed the same prompt. Two days earlier, the same feature had delivered a genuine throughput win on the same hardware. The field report, including the benchmark that looked healthy while the agent was broken, is vLLM MTP quietly breaking tool calls.

This is one dated configuration failure, not a claim that speculative decoding is generally unsafe. It was concurrency-dependent and did not reproduce in single-request testing, so a clean single-request test could have missed it. Pin the version or assert the disabled setting at startup; an upgrade can re-enable an optimization you intentionally removed.

Measure successful tool calls, not token speed

The load test needs an assertion for successful, well-formed tool invocations divided by attempts, at the concurrency the agents use. Otherwise a fast benchmark will report success on a broken lane.

If raw output is correct but the client mishandles it, fix the client’s tool-call handling. If inference comes from a hosted API, report the behaviour and switch models; you do not have the flags or visibility. If calls succeed but the agent chooses the wrong tool, the problem is model capability or prompt design, not this serving stack.

Bottom line

Start with raw output and speculative decoding, one change per restart. Then test unquantized weights and a matching parser/template before moving engines. Tool-call correctness is a separate metric from throughput.

Reference