In a single turn, the model emitted three tool calls: reserve_seat, charge_card, and send_confirmation. Your runtime read the response, saw three call blocks, and did what most LangChain-shaped glue code does: it ran them concurrently with an asyncio.gather. The seat was reserved. The confirmation email went out. The charge failed, because the card had expired and the vendor's 3DS flow timed out. On Monday morning the support queue had seven of these, and the finance team wanted to know why confirmation emails were being sent for transactions that never happened.

Function calling is taught as one call per turn. That was true in the early docs and demos. It stopped being true at least two years ago. Both Anthropic and OpenAI default to parallel tool use: a single assistant message can return a list of tool call blocks, each with its own tool_use_id or tool_call_id, and your runtime is expected to execute all of them and return one tool-results message with every result in it. If your runtime fires them all concurrently without thinking about ordering or safety, you get races, half-finished side effects, and the kind of partial-state bugs that no single-tool test suite catches.

Issue 19 argued that most of your agents should be workflows. Issue 16 covered the server side of MCP in production. This issue sits between the two: the small piece of runtime logic that decides, for a given model response, which of its tool calls can safely fan out and which must run in order. The deliverable is a tool execution planner that tags each tool as a parallel-safe read or an ordered write, runs reads concurrently and writes one at a time in the order emitted, and returns every result with its call ID in a single message so the model's next turn sees the full picture.

Why providers return multiple tool calls per turn

Parallel tool use is a latency optimisation on the vendor's side. When the model's planning step decides that it needs to look up the user, the account state, and the policy to answer the question, the model doesn't need to see any of those results before emitting the next call. All three are independent queries. Emitting them in one response saves two round-trips to the vendor, which saves a second or two of wall-clock time on a user-facing agent. OpenAI's parallel_tool_calls flag defaults to true, and Anthropic ships disable_parallel_tool_use so you can turn the behaviour off (see Further reading for both docs).

The vendors are right about reads. Three concurrent SELECT queries are safe to fan out. The problem is that the model doesn't know the difference between a read and a write at the schema layer, and nothing in the function-calling protocol stops it from emitting charge_card and send_confirmation in the same response. In the early days this was rare; the models that shipped with parallel tool use tended to serialise writes anyway because the training data had writes in sequence. By mid-2026 that stopped being reliable. The planner has to enforce the ordering, because the model won't.

There's a second issue that only shows up once you look at the runtime. The provider contract wants every result returned in one tool_results message, keyed by the call ID the model emitted. If one of your tool calls errors, you still have to return a result for it; otherwise the model's context goes missing a block and the next turn behaves oddly. "Fail the whole turn" is a common instinct for backend engineers and it's the wrong one here. Partial-failure reporting is table stakes.

The common wrong approach

The default pattern most teams ship first is "iterate the tool calls, fire them concurrently, collect the results". It looks like this in Python:

# The wrong default
async def execute_turn(response):
    tasks = [
        execute(call) for call in response.tool_calls
    ]
    results = await asyncio.gather(*tasks)
    return build_tool_results_message(results, response.tool_calls)

Three things are wrong with this. First, it treats every tool as parallel-safe, which is exactly the bug in the opening scene. Second, if one of the calls raises, the gather cancels the siblings or returns an exception back up the stack, and now you have one write that completed, one that was cancelled mid-flight, and one that never started. Third, the error path usually surfaces to the caller as a single failed turn, so the model never gets to see which specific call failed and which succeeded.

Each of those three failures is a different kind of bug in the eventual postmortem, and all three come from the same root: treating "the response has a list of tool calls" as if it means "the tool calls are independent".

The tool execution planner

The planner has five responsibilities. Keep it as a thin layer between your tool registry and the vendor-facing turn loop; it is the single place that gets to decide what runs in what order.

Tag every tool. At registration time, every tool gets a mode: read (parallel-safe, idempotent, no observable side effects), write (changes state the user or another system can observe, must be ordered), or exclusive (a write that must run alone on its turn, like a payment). Make the tag live in the same file as the tool handler so a developer adding a new tool has to pick one. Default should be write, not read; the safe choice is the conservative one.

Split the model's emitted calls into a plan. Walk the tool calls in the order the model emitted them. Group consecutive reads into a parallel batch. Writes go into the plan as singletons, in order. If an exclusive tool shows up, the plan collapses to a single call for that turn and the rest are deferred with a "skipped: write ordering" result.

Run reads concurrently, writes serially. For each read batch, fire the calls under asyncio.gather with return_exceptions=True so an error in one doesn't take down the siblings. For each write, await it; if it fails, do not advance to the next write, and tag the remaining writes "skipped: prior write failed" in the result message.

Return every result with its call ID. The vendor-facing message must contain exactly one result block per call ID the model emitted. If a call was skipped by the planner, the result block still exists; its content is a short structured error the model can read ("skipped: write ordering" or "skipped: prior write failed"). The model's next turn now has a full picture of what ran and what didn't, and can recover.

Report partial failure per call. The turn itself never fails. If every call errored, the message still gets sent back to the model with every error in it. The orchestrator above the planner decides whether to continue the agent loop or stop, based on which specific calls failed (that is Issue 13's bounds territory).

The planner reads the three tool calls, fans the two reads out in parallel, awaits them, then runs the write afterwards. The model sees one tool_results message with all three blocks keyed by their call IDs. The ordering is explicit on the server side, which means it survives model versions.

A worked example: the booking turn

The opening scene becomes safe with the planner in place. The model emits reserve_seat (write), charge_card (exclusive), and send_confirmation (write). The planner groups them:

  • Plan: reserve_seat as singleton write 1, charge_card as exclusive, send_confirmation deferred because exclusive collapses the turn.

The runtime executes reserve_seat first. If it succeeds, it runs charge_card. If the charge succeeds, the planner returns three result blocks: a success for reserve_seat, a success for charge_card, and a "skipped: exclusive-write collapsed this turn, please re-emit" for send_confirmation. The model's next turn re-emits send_confirmation with the IDs from the previous results, and the planner runs it as a standalone write.

If charge_card fails, the result block for it carries the structured error, and send_confirmation is returned "skipped: prior write failed". The model's next turn sees that the reservation went through but the charge didn't, and can choose to call release_seat as a compensating action. The confirmation email, critically, never fired.

The behaviour difference versus the naive runtime is small in code and enormous in outcome. One wrapper layer stopped the "email sent for a transaction that never happened" bug on every booking turn in your product.

Turn-level overrides

Sometimes the simplest fix is to turn parallel tool use off for a specific turn. The vendors support this. Anthropic's disable_parallel_tool_use: true on the request forces the model to emit one call per response. OpenAI's parallel_tool_calls: falsedoes the same thing. The trade is latency: you add a round-trip per tool call, so a three-tool turn becomes three turns instead of one.

Switch parallel calls off on turns that expose write tools, and leave it on when the available tools are all reads. If your agent exposes get_user, get_account_state, lookup_policy for the discovery phase, keep parallelism on. When the same agent moves to the action phase and the tool list includes reserve_seat and charge_card, switch it off. Splitting read and write phases also makes the agent's behaviour easier to trace, because the sequence in the log matches the sequence that actually happened.

The planner approach above is more work than the per-turn flag but survives more cases. A product-ready agent usually ends up with both: the flag for turns where the right thing is obvious, the planner for turns that mix safe reads and ordered writes.

Common mistakes

Firing everything with asyncio.gather. The default behaviour of every quickstart tutorial for parallel tool calls. It works until the first time a write collides with another write or with a read that depends on the write's outcome. Replace it with the planner.

Returning only the successful results. When one call fails, backend engineers' instinct is to fail the whole turn and send a 500 back up the stack. The vendor contract wants one result per call ID. If you omit a block, the model's next turn is operating on a context with a hole in it, and the behaviour gets strange. Return every result, including the structured errors and the "skipped" blocks.

Trusting the model to serialise writes. Early docs implied that models would emit writes one at a time. That's not a contract and it stopped being true in practice by mid-2026. The server-side planner is what enforces the ordering, not the model.

Debugging without call-ID logs. When a parallel tool turn goes wrong, you want to see which call IDs the model emitted, which block the planner put them in, which ran, which were skipped and why. Log all of that into the Issue 4 trace with the model's response ID as the correlation key. "The second tool call failed" is useless if you don't know which the second one was.

The takeaway

Parallel tool use is a vendor default that saves latency on reads and costs you reliability on writes if you treat the response's tool-call list as a flat set to gather over. The planner is the small, boring piece of runtime logic that reads the tags on your tools, groups reads into parallel batches, runs writes serially in the order emitted, falls back to "skipped: exclusive" when a payment-shaped tool shows up with siblings, and returns a result block for every call ID the model emitted so the next turn has the full picture. Partial failure gets reported per call, not per turn. Writes stop running in parallel with each other. The confirmation email stops going out before the charge lands.

Production checklist

  • Add a mode field to every tool registration (read, write, exclusive) in the same file as the handler. Default to write for new tools.

  • Build the planner as a single module between the tool registry and the turn loop. Make it the only place that decides execution order.

  • Group consecutive reads in the model's emitted order into parallel batches; run each batch under asyncio.gather(..., return_exceptions=True).

  • Run writes serially in the emitted order. On a write failure, do not advance to the next write; mark the remaining writes as skipped with a structured error.

  • Collapse the turn when an exclusive tool is in the plan: run only that one, defer every sibling with a "skipped" block the model can read and re-emit next turn.

  • Return exactly one result block per call ID the model emitted. Every error, including "skipped", gets a structured body the model can parse.

  • Set the Anthropic disable_parallel_tool_use flag (or OpenAI parallel_tool_calls: false) on turns that expose write tools when you want the simpler guarantee at the cost of a round-trip per call.

  • Log the model response ID, every call ID, the planner's grouping decision, and each call's outcome in the Issue 4 observability trace. Correlate on the response ID.

  • Add a parallel-call test to the Issue 3 golden set that emits a mix of reads and writes and asserts the planner's ordering decision, not just the final answer.

  • Review the plan on every model migration (Issue 14). Preference-tuning drift can change how often the model emits writes in the same response as reads.