AI incident runbooks

Ten runbooks for the incidents AI systems actually have.

Each one starts with evidence and containment in the first fifteen minutes and the first hour, then lists the evidence to collect, containment options, durable controls, communications, and closure criteria. Written by the people who run these incidents; published so you can run them without us.

Version 1.0 · 27 August 2026Free to use inside your organizationGeneric: adapt to your environment
Trigger. A key, token, service principal, delegated token, or provider credential associated with an AI service is suspected exposed or used outside its intended workload.
Severity guide. Treat as Severity 1 when use is active, high-cost models are reachable, cross-tenant authority exists, or the credential can create resources or secondary credentials.
Trigger. Unexpected token, request, model, GPU, or hosted-inference cost suggests unauthorized use, gray-market resale, proxying, or denial-of-wallet activity.
Severity guide. Severity 1 for rapidly increasing spend, capacity exhaustion, enterprise customer impact, or evidence of stolen identity. Severity 2 for bounded abuse without material service impact.
Trigger. A trusted MCP server, tool definition, OAuth scope, command, package, certificate, or endpoint changes unexpectedly, or an agent invokes a tool whose effective behavior no longer matches approval.
Severity guide. Severity 1 when a changed tool can access secrets, code, customer data, cloud control planes, or execute commands.
Trigger. An agent memory, profile, vector store, retrieval source, protected key, or durable context contains attacker-controlled content that can influence later privileged actions.
Severity guide. Severity 1 when poisoned state affects multiple users, persists across sessions, reaches privileged tools, or can cause data disclosure or execution.
Trigger. An agent reads, writes, executes, sends, publishes, transfers, or changes a resource outside intended authority or user intent.
Severity guide. Severity 1 for customer data, secrets, code, cloud control planes, destructive action, cross-tenant effect, or external disclosure.
Trigger. An undocumented or unapproved inference server, agent, MCP server, model gateway, notebook, vector store, Ray cluster, control panel, or AI integration is discovered internally or externally.
Severity guide. Severity 1 when internet-exposed, unauthenticated, connected to production data or credentials, or capable of code execution.
Trigger. A model file, pickle, dataset, agent skill, MCP package, plugin, coding extension, container, or dependency is suspected malicious, trojanized, untrusted, or replaced.
Severity guide. Severity 1 when loaded in production, capable of execution, connected to secrets or source code, or distributed to customers.
Trigger. Evidence suggests data, model, key, prompt, output, or workload state crossed a claimed private, local, sovereign, air-gapped, or confidential boundary.
Severity guide. Severity 1 for regulated, national, classified, or high-value customer data; attestation or key-custody failure; or systematic boundary bypass.
Trigger. A hosted AI provider, model gateway, proxy, marketplace, or external tool service reports or exhibits compromise, abuse, outage, unexpected behavior, model substitution, key exposure, or data-handling failure.
Severity guide. Severity 1 when customer data or prompts are exposed, provider authority is abused, model behavior changes materially, or business-critical service is unavailable.
Trigger. An AI coding agent or connected tool makes unauthorized code, dependency, workflow, secret, branch, release, or infrastructure changes, or consumes malicious repository content.
Severity guide. Severity 1 for production release, secret theft, signing-key access, CI control, protected-branch bypass, or customer distribution.

How to use these. They are generic by design: they do not know your environment, your provider, or your legal obligations. Free to use and adapt inside your organization; keep the source line if you republish. Version 1.0, 27 August 2026. Every runbook ends with the same rule: before closure, convert the incident into a diagram, a root-cause finding, a hunt analytic, a regression test, a validation memo, and an owner. That conversion is the part most teams skip, and it is the part we are hired for.

Before any adversarial work we sign a written scope. The rules-of-engagement template is public.

Happening now?

Send a project inquiry and set timing to active incident. Joey Victorino reads those first and answers the same business day, US Pacific. Put no credentials, prompts, or customer data in the form.

Describe the situation