Leaked AI or Model API Credential
A key, token, service principal, delegated token, or provider credential associated with an AI service is suspected exposed or used outside its intended workload.
Trigger: A key, token, service principal, delegated token, or provider credential associated with an AI service is suspected exposed or used outside its intended workload.
Severity guide: Treat as Severity 1 when use is active, high-cost models are reachable, cross-tenant authority exists, or the credential can create resources or secondary credentials.
First 15 minutes
- Open a case and assign incident commander and technical lead.
- Record the credential identifier, owner, intended workload, provider, entitlement, discovery source, and first known exposure time. Never place the secret itself in the ticket.
- Preserve gateway, provider, cloud identity, WAF, billing, and source-control logs before making broad changes.
- Constrain or disable the credential using the least destructive action that stops abuse.
- Start a provider-side usage and cost watch for all related identities and models.
First 60 minutes
- Build the credential lineage: where created, stored, copied, injected, and used.
- Search all provider, gateway, cloud, and network logs for use before and after the exposure window.
- Identify downstream sessions, proxies, service principals, resources, keys, or tokens created with the original authority.
- Create a replacement bound to a named workload identity. Do not distribute another shared secret.
- Sequence dependent rotations and validate legitimate workloads after each change.
- Hunt for continued behavior after rotation, especially from new keys, identities, origins, regions, or models.
Evidence checklist
- credential ID and redacted fingerprint
- owner and intended workload
- source locations, repositories, CI systems, and secret stores
- model and region entitlement
- origin IP, ASN, device, and geography
- requests, concurrency, token volume, latency, and cost
- secondary resources, sessions, or credentials created
- rotation, disablement, and provider-revocation timestamps
Containment options
- disable or scope the credential
- remove unauthorized entitlements
- block suspicious origins at gateway or provider where safe
- freeze high-risk model or region access
- apply emergency quota and spend ceilings
- revoke sessions and secondary authority
- preserve evidence before deleting attacker-created resources
Recovery and durable controls
- replace shared keys with workload identities or short-lived tokens
- centralize secret injection and rotation
- add repository and CI secret scanning
- baseline usage by workload and model
- test analytics with synthetic use
- document owner, entitlement, expiry, and emergency revocation path
Communications
- Security, AI platform, cloud, finance, and product are mandatory stakeholders.
- Notify legal or privacy if prompt, customer, or regulated content may have been exposed.
- Coordinate with the AI provider and existing IR partner through a named evidence route.
- Executive updates state known spend, known scope, containment state, and uncertainty.
Closure criteria
- Unauthorized use has stopped for an agreed observation window.
- All related authority is inventoried, rotated, or revoked.
- Legitimate workloads operate on named identities.
- Abuse analytics and spend controls are active and tested.
- Root cause, residual risk, and owners are accepted.
Required conversion to practice
Before closure, produce:
- one updated attack-path or trajectory diagram
- one root-cause finding
- one production or candidate hunt analytic
- one safe replay or regression test
- one remediation-validation memo
- one owner and residual-risk decision
How to use these. They are generic by design: they do not know your environment, your provider, or your legal obligations. Free to use and adapt inside your organization; keep the source line if you republish. Version 1.0, 27 August 2026. Every runbook ends with the same rule: before closure, convert the incident into a diagram, a root-cause finding, a hunt analytic, a regression test, a validation memo, and an owner. That conversion is the part most teams skip, and it is the part we are hired for.
- Inference Cost Spike or Token-Jacking
- MCP Tool Poisoning, Impersonation, or Rug Pull
- Persistent Memory or RAG Poisoning
- Unauthorized Agent Action or Data Exfiltration
- Shadow or Exposed AI Infrastructure
- Malicious Model, Artifact, Skill, Plugin, or Extension
- Confidential or Sovereign AI Boundary Failure
- AI Provider or Model Gateway Third-Party Incident
- AI Coding Agent Repository or CI Compromise
Happening now?
Send a project inquiry and set timing to active incident. Joey Victorino reads those first and answers the same business day, US Pacific. Put no credentials, prompts, or customer data in the form.