Runbook RB-06

Shadow or Exposed AI Infrastructure

An undocumented or unapproved inference server, agent, MCP server, model gateway, notebook, vector store, Ray cluster, control panel, or AI integration is discovered internally or externally.

Version 1.0 · 27 August 2026Free to use inside your organization

Trigger: An undocumented or unapproved inference server, agent, MCP server, model gateway, notebook, vector store, Ray cluster, control panel, or AI integration is discovered internally or externally.

Severity guide: Severity 1 when internet-exposed, unauthenticated, connected to production data or credentials, or capable of code execution.

First 15 minutes

  • Confirm the asset is in authorized scope and identify owner, environment, and network path.
  • Capture passive evidence and service fingerprint before active validation.
  • Apply emergency network restriction for clearly dangerous exposure with owner approval.
  • Preserve cloud, DNS, certificate, endpoint, and deployment evidence.
  • Open an ownership and data-classification workstream.

First 60 minutes

  • Identify process, package, image, model, agent, tools, credentials, and data sources.
  • Determine exposure, authentication, tenant, and egress controls.
  • Search cloud, code, endpoint, and DNS evidence for sibling assets.
  • Review access logs for historical use and suspicious origins.
  • Establish whether the asset is development, test, abandoned, or production.
  • Create disposition: approve and onboard, isolate, or retire.

Evidence checklist

  • IP, host, DNS, certificate, cloud resource, and endpoint
  • service and version fingerprint
  • owner, environment, purpose, and deployment source
  • authentication and authorization
  • model, data, tools, credentials, and network reachability
  • historical access and resource use
  • CMDB and ticket state

Containment options

  • restrict network exposure
  • disable default or shared credentials
  • remove production secrets and data
  • isolate or stop unauthorized service
  • revoke external routes and certificates
  • onboard approved asset into logging and inventory

Recovery and durable controls

  • repeatable multi-source discovery
  • AI asset inventory fields and owner SLA
  • deployment policy and approved patterns
  • default logging, identity, quota, and egress controls
  • drift detection and periodic exposure review

Communications

  • Work with engineering to avoid treating shadow infrastructure as misconduct before facts are known.
  • Escalate unowned critical assets to the platform executive.
  • Notify legal only when data or customer exposure is plausible.

Closure criteria

  • Asset is onboarded, isolated, or retired.
  • Historical use is reviewed or bounded.
  • Owner and data classification are recorded.
  • Discovery method is rerun and documented.
  • Preventive deployment controls are assigned.

Required conversion to practice

Before closure, produce:

  • one updated attack-path or trajectory diagram
  • one root-cause finding
  • one production or candidate hunt analytic
  • one safe replay or regression test
  • one remediation-validation memo
  • one owner and residual-risk decision

How to use these. They are generic by design: they do not know your environment, your provider, or your legal obligations. Free to use and adapt inside your organization; keep the source line if you republish. Version 1.0, 27 August 2026. Every runbook ends with the same rule: before closure, convert the incident into a diagram, a root-cause finding, a hunt analytic, a regression test, a validation memo, and an owner. That conversion is the part most teams skip, and it is the part we are hired for.

Happening now?

Send a project inquiry and set timing to active incident. Joey Victorino reads those first and answers the same business day, US Pacific. Put no credentials, prompts, or customer data in the form.

Describe the situation