Search

Why Your LLM Spend Limit Doesn't Actually Stop Spending

Spend limits read stale balances, so they stop nothing under concurrency. Test with xargs -P 20, read the 200/429 difference, and fix it with a gateway cap.

Mohit7 min read

Filed under Agent Reliability· see every report on this topic

Four concurrent requests each pass the same stale budget check and collectively cross the cap.

Verdict. Most budget systems check spend before a request and record it after, without a lock between those actions. Concurrent requests read the same total, pass the check, and run. The cap uses stale numbers, so it is a dashboard rather than a control. Set a hard provider-console limit first, then a per-key gateway cap. Use reserve-before-execute when you need per-customer attribution.

The operator and the leak

You run agents or a gateway that bills tenants, and a cost cap has to stop money leaving the account. A cap that holds at concurrency one and fails at concurrency ten is a race, not a bad setting.

The failure sequence

The broken order is check → execute → record. One request at a time works. Two overlapping requests do not, because nothing holds a lock across those three actions.

Both requests read the same pre-request total and compare it to the same limit. Both pass and execute. Each check was correct at that instant, but the total was stale before the request mattered.

This is a check-then-act race applied to money. Two requests in the same window are enough. Agent workloads create that pattern because an agent chooses its own request volume.

Evidence across four layers

The reports come from four layers with separate implementations. Every title below was fetched and confirmed verbatim on 2026-08-22. These are user reports and open issues, not confirmed vendor behaviour.

Layer Issue What it reports
Model vendor CLI #83048 budget.spent() reports 72x under actual consumption, filed SEV-1
Model vendor CLI #85022 Drained a prepaid balance without consent
Gateway #18730 Concurrent requests bypass TPM limits, ~6x over a 100 TPM cap
Gateway #26672 A key with max_budget: 0.05 reached spend: 0.540408 and kept serving
Gateway #35524 reserve_budget_for_request() returns without reserving when cost cannot be estimated
Sandbox provider daytona #4791 A 2 vCPU sandbox burned ~$5 of credit in one hour
Protocol MCP #3229 RFC asks for token metering and session budgets because the protocol has none

Four related reports sit behind these: Claude Code #77095, #77373, #85400, and LiteLLM #12905 on unenforced team-key budgets.

Read #35524 twice. The reservation mechanism exists, but the report says it does nothing when a request’s maximum cost is hardest to estimate. It falls back to read-time enforcement instead of admission control.

Test your setup

Run two tests in about twenty minutes. First, check whether the meter is honest. Then check whether the cap holds.

Check meter drift

Record your platform’s reported spend and the provider billing-console figure for the same 24-hour window, then divide.

# LiteLLM example: spend recorded for a key over the last day.
# Prerequisite: LITELLM_URL and LITELLM_MASTER_KEY set; key_name is the key you are auditing.
curl -s "$LITELLM_URL/spend/logs?start_date=2026-08-21&end_date=2026-08-22" \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
| python3 -c 'import json,sys; print(round(sum(r["spend"] for r in json.load(sys.stdin)), 6))'

Expected output is one number. Compare it with the same window in the provider billing console. A ratio near 1.0 means the meter agrees; anything above 1.05 needs an explanation before you trust a cap based on it.

Force the race

A sequential test passes on a broken system. Overlap the requests instead.

# Fire 20 concurrent requests against a key with a deliberately small budget.
# Prerequisite: TEST_KEY has a max_budget you are willing to lose (e.g. $0.05).
seq 20 | xargs -P 20 -I{} curl -s -o /dev/null -w '%{http_code}\n' \
  "$LITELLM_URL/v1/chat/completions" \
  -H "Authorization: Bearer $TEST_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hi"}],"max_tokens":16}' \
| sort | uniq -c

-P 20 runs requests in parallel. Expect a mix of 200 and 429, then read the key’s final recorded spend and compare it with the budget. If final spend exceeds the budget, the cap is advisory. If every request returns 200, it is not enforcing at all.

Fixes, cheapest first

Hard limit at the provider

Set a billing limit in the provider console. The party holding the money enforces it, it needs no code, and it covers the catastrophic case. Most providers apply it with delay, so use it as a backstop rather than a precise cap.

My take: for a solo builder with one key, stop here.

Per-key gateway caps and alerting

Use a gateway cap to isolate tenants. It can still overshoot under concurrency, per #18730 and #26672. Alert on the drift ratio from the test so you learn about it in minutes instead of at invoice time.

Reserve before execute

Change the order so the wallet is debited before the request runs:

authenticate → authorize → reserve → execute → meter → settle

reserve atomically debits an estimated maximum, so a second concurrent request sees the first reservation and is refused admission. settle returns the unused difference and writes both entries to an append-only ledger.

Two conditions matter. The reservation must be atomic. A read followed by a write recreates the race. Estimation must fail closed. If maximum cost cannot be estimated, refuse the request or charge a conservative ceiling. #35524 reports the opposite path.

The reporters propose this fix. #18730’s submitter names AWS Bedrock’s token-reservation pattern; MCP #3229 asks the protocol to carry metering because server operators absorb inference costs without a standard per-transaction billing path. BricksLLM, Bifrost, agentgateway, and any-llm-gateway each built a proxy around stolen keys, missing per-key metrics, and config sprawl.

KitDev’s gateway uses this sequence with an append-only ledger, while Hermes and Pi run inside Firecracker sandboxes. These are development proofs on self-hosted infrastructure, not a production availability claim. Cost-per-request figures from that setup are to be captured from a live test.

What this does not fix

Estimation error sets the ceiling. A maximum reservation must be pessimistic. Long-context requests reserve large amounts and release most at settlement, so a customer near the limit can be refused capacity they will not use.

The ledger is on the critical path. Atomic reservation writes before every request. Decide whether a wallet-store outage refuses traffic or disables enforcement.

Non-token meters stay out of view. Sandbox time, egress, and storage use separate meters. Daytona #4791 is compute-time burn; a gateway that only meters model calls cannot see it.

Enforcement is separate from attribution. A global cap can hold perfectly while you still cannot identify the customer who spent it. Attribution requires tenant identity in the ledger, which changes the data model rather than a limit setting.

When it is something else

If the numbers are wrong but the cap holds in the concurrency test, you have a reporting problem rather than an enforcement problem. Start with the drift ratio and the provider billing export, not gateway configuration.

If spend is high but every cap works, inspect workload design. Retry storms and context resends are usual causes. A tool-call failure can silently trigger retries; see fixing broken tool calls in self-hosted inference.

Reference