Tool Call Orchestration: Sequential, Parallel, and DAG Execution
Last Updated on August 19, 2026 by Editorial Team
Author(s): Armin Norouzi, Ph.D
Originally published on Towards AI.
Tool Call Orchestration: Sequential, Parallel, and DAG Execution
A research agent calling 6 tools sequentially waits 1,832ms at P50. The same 6 tools run in parallel — ignoring all dependencies — finish in 671ms. With a dependency-aware DAG that respects which tools must complete before others can start, the critical path lands at 1,358ms, 1.35× faster than sequential while honoring correctness constraints. The difference between these three orchestration strategies is not a minor optimization; it determines whether your agent fits within a user-facing latency budget.

The article explains how orchestration strategy affects latency in multi-tool research agents: sequential simply sums tool latencies, naive parallelism (ignoring dependencies) can be fast but incorrect when later tools rely on earlier outputs, and DAG-based execution respects dependencies while still extracting parallelism within dependency “waves.” It then analyzes latency distributions (especially tail behavior like P95/P99) using simulated trials, showing that DAG-based execution often beats sequential across percentiles with a smaller spread because variance doesn’t always stack when independent tasks overlap. The author uses critical-path reasoning to identify which tools dominate end-to-end latency and discusses how dependency “tax” explains the gap between parallel and DAG results. Additional sections cover retry overhead and how retries can disproportionately harm tail latencies under DAG constraints, provide selection rules for when to use sequential vs parallel vs DAG, and outline production implementation practices (explicit dependency declarations, per-wave timeouts, and returning partial results). Finally, it validates the model with asyncio-based benchmarking, demonstrates integrating DAG orchestration with LLM tool-use APIs, and summarizes the methodology and assumptions behind the reported latency and failure-rate experiments.
Read the full blog for free on Medium.
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.