Inference Cost Spike or Token-Jacking
Unexpected token, request, model, GPU, or hosted-inference cost suggests unauthorized use, gray-market resale, proxying, or denial-of-wallet activity.
Trigger: Unexpected token, request, model, GPU, or hosted-inference cost suggests unauthorized use, gray-market resale, proxying, or denial-of-wallet activity.
Severity guide: Severity 1 for rapidly increasing spend, capacity exhaustion, enterprise customer impact, or evidence of stolen identity. Severity 2 for bounded abuse without material service impact.
First 15 minutes
- Confirm the cost signal against provider and billing data.
- Separate customer growth, test traffic, retry storms, and product defects from suspected abuse.
- Identify affected provider accounts, models, regions, deployments, and workload identities.
- Apply reversible cost containment: quota, concurrency, model, region, or origin restriction.
- Preserve raw request and billing detail.
First 60 minutes
- Baseline the identity against prior 7, 30, and 90-day model, origin, concurrency, token, and time-of-day behavior.
- Look for fan-out across many IPs or users behind one credential, reverse-proxy signatures, and rapid model substitution.
- Correlate sign-ins, key creation, source-control exposure, and cloud-resource creation.
- Estimate current burn rate and worst-case exposure.
- Engage provider abuse and billing teams with timestamps and request identifiers.
- Start an adaptation watch after containment.
Evidence checklist
- provider account and deployment
- identity or credential identifier
- model, region, origin, and user-agent
- requests, input/output tokens, latency, concurrency, and cost
- quota and rate-control decisions
- proxy and gateway headers
- sign-in and key-creation events
- customer and workload baselines
Containment options
- reduce entitlement to required models and regions
- cap spend, tokens, concurrency, and resource creation
- disable compromised identities or keys
- block abusive reverse-proxy paths
- isolate self-hosted endpoints
- protect legitimate priority workloads with separate identities and budgets
Recovery and durable controls
- per-workload identity and chargeback
- behavioral usage baselines
- automated spend guardrails and alerting
- provider and gateway log retention
- abuse classification and response ownership
- quarterly simulation of token-jacking scenarios
Communications
- Finance receives burn-rate and avoided-cost estimates.
- Product receives customer-impact and capacity information.
- Legal and customer trust are engaged if tenant data or prompts may have transited attacker infrastructure.
- Do not claim compromise when evidence supports only product abuse.
Closure criteria
- Spend and capacity return to expected range.
- Abusive identities, origins, and proxy paths are contained.
- All usage is attributable to an approved workload or customer.
- Detection catches a synthetic replay.
- Business owners accept residual abuse economics.
Required conversion to practice
Before closure, produce:
- one updated attack-path or trajectory diagram
- one root-cause finding
- one production or candidate hunt analytic
- one safe replay or regression test
- one remediation-validation memo
- one owner and residual-risk decision
How to use these. They are generic by design: they do not know your environment, your provider, or your legal obligations. Free to use and adapt inside your organization; keep the source line if you republish. Version 1.0, 27 August 2026. Every runbook ends with the same rule: before closure, convert the incident into a diagram, a root-cause finding, a hunt analytic, a regression test, a validation memo, and an owner. That conversion is the part most teams skip, and it is the part we are hired for.
- Leaked AI or Model API Credential
- MCP Tool Poisoning, Impersonation, or Rug Pull
- Persistent Memory or RAG Poisoning
- Unauthorized Agent Action or Data Exfiltration
- Shadow or Exposed AI Infrastructure
- Malicious Model, Artifact, Skill, Plugin, or Extension
- Confidential or Sovereign AI Boundary Failure
- AI Provider or Model Gateway Third-Party Incident
- AI Coding Agent Repository or CI Compromise
Happening now?
Send a project inquiry and set timing to active incident. Joey Victorino reads those first and answers the same business day, US Pacific. Put no credentials, prompts, or customer data in the form.