Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
GitHub Copilot Dynamic Workflows: Build Incident Response Agents You Can Debug
Latest   Machine Learning

GitHub Copilot Dynamic Workflows: Build Incident Response Agents You Can Debug

Last Updated on October 6, 2026 by Editorial Team

Author(s): Ethan Mark

Originally published on Towards AI.

GitHub Copilot Dynamic Workflows: Build Incident Response Agents You Can Debug

GitHub Copilot Dynamic Workflows: Build Incident Response Agents You Can Debug

A practical implementation guide for turning one frantic, open-ended incident prompt into a bounded investigation your team can inspect, pause, and trust.

More agents help only when their handoffs and evidence stay visible.

At 2:13 a.m., “look into the outage” is a terrible agent prompt. It asks a system to choose its own scope, tools, stopping point, and definition of proof while the people on call need exactly the opposite: a short timeline, a clear list of unknowns, and no surprise production changes.

GitHub Copilot’s new dynamic workflows create a useful middle path. They let you put the orchestration in code while using agents for the parts that genuinely benefit from judgment. A workflow can run deterministic collection steps, fan out independent analysis, validate the results, pause for a human, and package a final handoff. GitHub even uses service-incident investigation as a first-class example.

The novelty is not “several agents in parallel.” Teams have been doing that for a while. The practical shift is that the workflow author defines the stages, conditions, and handoffs, rather than hoping an agent invents a sensible process during a stressful run. This guide shows how to apply that idea to a read-only incident-investigation workflow that is useful before you grant any agent write access.

Why a dynamic workflow is different from a fleet of agents

Copilot has modes for autonomy and for parallel subagents. Those are good for exploratory work. But an incident is a repeatable operational process. You want the same evidence sources, time window, stopping rules, and approval boundary every time, even if the agent’s diagnosis changes.

A dynamic workflow is a program in a Copilot extension. It can combine regular code, tools, agents, API calls, and user interaction. The stages may be sequential, parallel, or mixed. It can also ask agents for a specified result format, pause for review, and resume later. In Copilot CLI, you can monitor phases, active subagents, and AI-credit use, then pause or cancel a run; the Copilot app offers local-run monitoring too. The official usage guide also describes how to share an extension from a personal or repository extension directory.

That makes the workflow a small distributed system. Treat it like one. Give it contracts, timeouts, idempotent deterministic steps, durable state, and a way to explain why it stopped.

The useful mental model: deterministic code owns collection, permissions, validation, and state transitions. Agents own synthesis, anomaly interpretation, and clearly bounded hypotheses. Humans own production-changing decisions.

Choose a narrow first job: investigation, not remediation

Start with an incident workflow that can read data and draft a packet. Do not begin with a workflow that scales a cluster, disables an account, rolls back a deploy, or edits a routing rule. You will get fast feedback on quality without turning an uncertain model output into an irreversible action.

A strong v1 has one input: an alert or incident ID. Its output has four things: a time-bounded timeline, sourced observations, ranked hypotheses, and recommended next checks. It may create a draft ticket or a pull request only if your existing controls already make that output safe and reviewable.

This matches a recurring practitioner concern. Recent SRE discussions show broad interest in AI-assisted log triage and timeline drafting, but much less appetite for direct production writes. That contrast is your design opportunity: make investigation faster while keeping responsibility clear.

Define the run contract before you define an agent

The run contract is the answer to “what is this workflow allowed to know and do?” Put it in a typed input object and persist it with the run. It prevents an innocent request like “check payments” from expanding into an unbounded search across customer data.

  • Identity: incident ID, service, environment, and named incident commander.
  • Time box: a start and end time, plus a maximum lookback window.
  • Read lanes: the exact telemetry, deployment, ticket, and runbook tools the workflow may call.
  • Effect policy: read-only by default; any write requires a distinct approved action outside the investigation lane.
  • Budget: maximum elapsed time, agent count, tool calls, and AI credits.
  • Exit states: completed, needs-human-input, evidence-insufficient, policy-blocked, or failed.

Do not hide these rules in a prose prompt. A prompt can explain the purpose; code should enforce the envelope. If a tool request is outside the contract, deny it and record that denial as evidence. A blocked call is often useful information during an incident.

Split the investigation into independent evidence lanes

The best parallelism is not “ask three agents the same question.” It is a set of independent lanes that return comparable evidence. One lane cannot overwrite another lane’s conclusions. Each returns observations first, then a limited interpretation.

A bounded investigation flows from evidence collection to a human-reviewed packet.

For a web-service incident, three lanes are enough:

  • Telemetry lane: fetch error-rate, latency, saturation, and trace exemplars for the approved time window.
  • Change lane: collect deploys, feature flags, configuration changes, and relevant pull requests near the onset time.
  • Knowledge lane: retrieve service ownership, known failure modes, runbooks, and closely matching past incidents.

Run deterministic collection before the agents start. Then hand each lane a compact evidence bundle. This matters for cost and reliability: the agent should analyze a curated snapshot, not independently wander through production tools while the incident evolves.

GitHub’s guidance on multi-agent engineering makes the same case in more general terms: agents behave like distributed-system components, so shared state, ordering, and implicit handoffs become failure surfaces. Typed schemas and explicit actions turn “inspect logs and guess” into a specific contract failure you can repair or escalate.

Build one boring, reliable reference flow

For the first version, resist the urge to create a clever supervisor agent. Make the path obvious enough to draw on a whiteboard. First, validate the incident ID and resolve the approved service and time window. Second, run the deterministic collectors with read-only credentials and save their raw receipts. Third, launch the three evidence lanes against those frozen bundles. Fourth, reject incomplete or unsupported lane outputs. Fifth, let one synthesis step create a short packet from the accepted records. Sixth, pause for the incident commander.

That order has a few quiet advantages. It avoids a race where one agent sees a new deploy while another does not. It keeps raw collection separate from model interpretation. It also means a rerun can reuse the same evidence snapshot if you need to compare models, prompts, or workflow versions. A workflow that is easy to replay is easier to improve.

Give every stage a small, honest failure mode. If the metrics query times out, record telemetry-unavailable; do not let the narrative agent quietly fill the gap. If the change lane has no access to a configuration system, state that in the final packet. If two agents disagree, preserve both claims and route the conflict to the human checkpoint. “No supported conclusion” is safer than a smoothly written fiction.

For tool calls that might be retried, use an idempotency key tied to the run ID and stage name. For external reads with mutable results, store the retrieval timestamp and query shape. For budgets, fail closed: when a lane hits its allotted calls or time, it returns an incomplete-evidence state rather than borrowing unlimited work from the rest of the run. These are ordinary workflow-engineering techniques. They matter more, not less, when some stages contain probabilistic reasoning.

Make every agent boundary machine-checkable

Natural-language summaries are pleasant to read and painful to compose. An agent might call a time range “last hour,” omit the source, or merge a fact with a guess. Downstream code should not have to infer which is which.

Write on Medium

Require a small schema at every handoff. Each observation needs a source reference, time range, confidence, and a statement that can be checked later. Each hypothesis needs supporting observation IDs and a disconfirmation test. The final narrator can turn that structured record into readable prose.

// Conceptual TypeScript: validate before a finding can flow onward.
type Observation = {
id: string;
lane: "telemetry" | "change" | "knowledge";
statement: string;
sourceRef: string;
observedAt: string;
confidence: "high" | "medium" | "low";
};
type Hypothesis = {
claim: string;
evidenceIds: string[];
disconfirmingCheck: string;
recommendedAction: "collect-more" | "page-owner" | "request-approval";
};
function acceptFinding(finding: Hypothesis, observations: Observation[]) {
const supported = finding.evidenceIds.every(id =>
observations.some(observation => observation.id === id)
);
if (!supported) throw new Error("Unsupported hypothesis");
return finding;
}

This is intentionally not a copy-paste SDK implementation. It is the important contract: no source, no claim; no disconfirmation test, no recommendation; no approved action type, no executable outcome. A schema also gives you a clean retry boundary. You can ask an agent to repair malformed output once, then escalate instead of letting it improvise forever.

Use checkpoints to keep human approval meaningful

A human approval button at the very end is not enough. By then, an agent may have gathered far more data than intended or spent the whole budget on a bad branch. Put a checkpoint after evidence collection and before any action proposal that could cause an external effect.

At the first checkpoint, show the incident commander the requested scope, sources consulted, missing permissions, and early conflict signals. Let them narrow the time window, add one approved source, or stop the run. At the second checkpoint, show only evidence-backed hypotheses and the proposed next action. The approval should name a person and expire; it should not be a vague “looks good.”

Fast incident response is not the same as automatic incident response. The fastest trustworthy workflow removes low-value searching, then hands a person a short decision with evidence attached.

Dynamic workflows are designed to pause and resume, which makes these checkpoints operational instead of ceremonial. Keep the durable run state outside a chat transcript: contract version, input snapshot, tool receipts, schema-valid findings, approval record, and workflow version. A later reviewer must be able to distinguish what was observed from what was generated.

Trace the workflow like an application, not a conversation

When a result is wrong, “the agent got confused” is not a diagnosis. You need to know whether collection failed, a tool returned stale data, a lane exceeded its budget, an agent violated a schema, or the final synthesis overreached.

Give every run a correlation ID. Attach it to deterministic calls, agent phases, tool calls, validation failures, approvals, and the final packet. Track at least:

  • run duration and time per lane;
  • tool-call count, denial count, and retry count;
  • schema repair attempts and unsupported-claim rejections;
  • AI-credit or token usage by phase;
  • human edits to the final timeline or hypothesis list; and
  • the eventual incident disposition, when available.

Copilot SDK supports OpenTelemetry and W3C trace-context propagation, so application spans can sit under CLI work and tool-handler spans can attach to the same distributed trace. That link is more valuable than a wall of chatbot text: it connects a generated conclusion to the collection and validation steps that preceded it.

Build the evidence packet people will actually use

Do not end with a generic “root cause analysis.” During an active incident, the reader needs a packet that answers the next operational question in under a minute.

  • Scope: service, environment, time window, workflow version, and current status.
  • Timeline: timestamped observations with source links or immutable references.
  • What changed: approved change-lane evidence, separated from coincidence.
  • Hypotheses: ranked, each with evidence, uncertainty, and the quickest disconfirming check.
  • Open questions: what the workflow could not access or verify.
  • Decision request: the narrowest human decision needed next.

The packet should state “no conclusion” when evidence is thin. That is a high-quality outcome. A workflow earns trust by making uncertainty visible, not by producing an answer for every alert.

The endpoint is a better human decision, not an autonomous production change.

A practical rollout plan

Use the first few runs as an evaluation program, not a launch. Replay past low-risk incidents with frozen evidence. Ask incident responders to score the packet: Did it identify the right time window? Did every claim point to evidence? Did it surface a useful next check? Did it omit anything a responder had to hunt for?

Then run shadow mode on live alerts. The workflow can prepare a packet without paging anyone or touching any production system. Compare it with the human-led timeline. Only after it consistently produces useful, bounded investigation work should you add a low-risk external action, such as creating a draft incident note. Keep mutable infrastructure operations in a separate workflow with stricter authorization and rollback checks.

Dynamic workflows are currently a public preview, so isolate vendor-specific calls behind a small adapter and pin the workflow’s input/output contract. The new runtime should make your process more testable, not turn your incident protocol into a moving target.

FAQ

What are GitHub Copilot dynamic workflows?

They are code-defined orchestrations in Copilot extensions that can mix deterministic steps, agents, tools, APIs, and human interaction. The workflow author controls stages, handoffs, conditions, and limits.

When should I use a dynamic workflow instead of Copilot fleet?

Use a dynamic workflow when the process must be repeatable and inspectable: incident investigation, release checks, batch review, or a long run that needs checkpoints. Fleet is better for open-ended parallel exploration where Copilot can decide the task split.

Can a dynamic workflow make production changes?

It can call tools, but production writes should not be the first use case. Start read-only, require an explicit approval boundary for an effect, and keep reversible, narrowly scoped mutations in a separate contract.

How do I keep agent findings from becoming unsupported claims?

Use a schema that requires source references and evidence IDs for every claim. Reject or repair outputs that cannot point to collected evidence, and require a disconfirming check for each hypothesis.

How should I monitor Copilot dynamic workflows?

Track run phases, active subagents, budget use, schema failures, tool denials, retries, and human edits. Send the same correlation ID into your telemetry system so deterministic calls and agent activity share one trace.

Are dynamic workflows ready for production?

GitHub documents the feature as public preview and subject to change. Treat it as an integration to evaluate: test with frozen incidents, run in shadow mode, version contracts, and keep your human and policy gates outside model prose.

Sources: GitHub Changelog, GitHub Docs: Dynamic workflows, GitHub Docs: OpenTelemetry instrumentation, and GitHub’s multi-agent engineering guide.

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.