SOP — HEROKU_API_KEY Drift Recovery
Owner: Operator (Kristerpher) + agent Last updated: 2026-05-03 First incident: 2026-05-03 03:48 UTC (during MBT v1 polish-sprint deploys) Related issues: #925, #891, #943
If the rotation handler raised
OldTokenInvalidError("Heroku rejected the rolling token before mint"), that is the typed signal added in #943. The dyno'sHEROKU_API_KEYis the drifted copy — re-sync from vault per Path A or B below.
What "drift" means here
Updated 2026-08-29 (#4544): GitHub Actions deploys were retired 2026-07-10
(ADR-0135/ADR-0136) — there is no longer a GHA repo secret in this picture.
Woodpecker CI reads HEROKU_API_KEY live from vault on every pipeline run
(VAULT_SECRETS="HEROKU_API_KEY" in .woodpecker/*.yaml, both staging AND
prod deploy pipelines as of #4544) — there is no static WP copy either. The
copies that must agree are:
- Vault —
/MooseQuest/heroku/HEROKU_API_KEYin Infisical. Source of truth. - Heroku config var —
HEROKU_API_KEYon each of the four Heroku apps (raxx-console-prod,raxx-console-staging,raxx-api-prod,raxx-api-staging). Read by the running app for vendor calls. - Session bootstrap —
HEROKU_API_KEYexported bysession-bootstrap.shfor agent sessions.
Drift = one of those values stops matching the others. Before #4544 the most-painful failure mode was "GH Actions secret is stale, every Heroku deploy fails" — that class of failure no longer exists because CI has no static copy to go stale. The remaining failure mode is Woodpecker's own run-time vault fetch failing or the fetched value not matching the live Heroku config var (see Symptom row below).
Symptom — deploy-prod fails at deploy-raptor-prod with a diagnosable HTTP code (#4544)
ERROR: RELEASE_VERSION PATCH failed HTTP 401: {"id":"unauthorized","message":"..."}
Post-#4544, the RELEASE_VERSION config-var PATCH in deploy-prod.yaml prints
the HTTP status code and (truncated) response body on any non-2xx response
instead of swallowing it. HTTP 401 here means the authorization itself is
invalid, not merely that it hasn't been copied somewhere — the same key the
pipeline loaded from vault is also what raxx-api-prod's config var and the
other 3 Heroku apps would receive, so propagating it further does not fix
anything and only spreads a bad key. Diagnose first, then act:
- Run the "Confirm the source of truth is valid" check below against the
vault value. If it returns
401→ the authorization itself is revoked or expired; mint a new Heroku authorization perheroku-api-key-rotation.md(vault write, then propagate to the 4 Heroku app config vars, then revoke the old authorization) — do NOT just re-propagate the same (already-invalid) vault value. - If it returns
200→ the vault value is valid; the pipeline's own vault fetch failed (auth to vault.raxx.app, CF Access, or a transient error) — checkwpc_load_vault_secrets.py's stderr output in the failed step and the Infisical machine-identity health, not the Heroku authorization.
This is precisely the failure mode that broke deploy-prod pipelines 7000
and 7800 for 5 weeks — see
docs/incidents/2026-08-29-deploy-prod-stale-wp-secret.md.
How to detect drift
Symptom 1 — git push heroku fails in CI
Deploy to Heroku via heroku CLI + credential helper (subtree split):
Error: The token provided to HEROKU_API_KEY is invalid. Please
double-check that you have the correct token, or run `heroku login`
without HEROKU_API_KEY set.
Symptom 2 — heroku run fails locally (different problem; usually means your local CLI auth is stale, not vault drift)
Confirm the source of truth is valid
heroku run --app raxx-console-prod --no-tty 'python -c "
import requests
from app.services import vault
t = vault.get_secret_value(\"HEROKU_API_KEY\")
r = requests.get(\"https://api.heroku.com/account\",
headers={\"Authorization\": f\"Bearer {t}\",
\"Accept\": \"application/vnd.heroku+json; version=3\"},
timeout=10)
print(\"vault token validates:\", r.status_code, \"OK\" if r.ok else r.json())
"'
If this returns 200 OK → vault is valid; the drift is in GH or Heroku config. Continue below.
If this returns 401 → vault itself is stale; you need to mint a new token via the rotate-from-console UI (#885). Stop here; that's a different runbook.
Recovery — if vault is valid, GH is stale
Historical (#4544): this section predates the 2026-07-10 GHA retirement
(ADR-0135/ADR-0136). Woodpecker CI has no GH Actions secret and no static WP
copy to go stale — it re-fetches vault on every pipeline run, so a
vault-is-valid/CI-is-stale split is no longer possible for the CI path. This
recovery path is retained for the (now much narrower) case where some other
GH-Actions-era consumer of this secret still exists; if you hit this section
troubleshooting a Woodpecker deploy failure, you are almost certainly looking
at the wrong runbook section — see the deploy-prod symptom row above
instead.
Path A — Manual paste (60 seconds)
# Read vault value (safe; runs in dyno, value goes only to your terminal)
heroku run --app raxx-console-prod --no-tty 'python -c "
from app.services import vault
print(vault.get_secret_value(\"HEROKU_API_KEY\"))
"' | tail -1
Copy the value, then:
- Open
https://github.com/raxx-app/TradeMasterAPI/settings/secrets/actions - Click
HEROKU_API_KEY→ "Update secret" - Paste the value
- Click "Update secret"
Re-run the failed deploy:
gh workflow run deploy-console.yml -f environment=production -f ref=main
Path B — Automated, requires GITHUB_API_SECRETS_TOKEN (#925)
Once GITHUB_API_SECRETS_TOKEN is in vault with secrets:write scope:
heroku run --app raxx-console-prod --no-tty 'python /dev/stdin' <<'PY'
import base64, os, sys, requests
from nacl.public import PublicKey, SealedBox
from app.services import vault
token = vault.get_secret_value("HEROKU_API_KEY")
gh = vault.get_secret_value("GITHUB_API_SECRETS_TOKEN")
H = {"Authorization": f"Bearer {gh}",
"Accept": "application/vnd.github+json",
"X-GitHub-Api-Version": "2022-11-28"}
r = requests.get("https://api.github.com/repos/raxx-app/TradeMasterAPI/actions/secrets/public-key", headers=H, timeout=15)
r.raise_for_status()
pk = r.json()
sealed = SealedBox(PublicKey(base64.b64decode(pk["key"])))
encrypted = base64.b64encode(sealed.encrypt(token.encode("utf-8"))).decode("ascii")
r = requests.put("https://api.github.com/repos/raxx-app/TradeMasterAPI/actions/secrets/HEROKU_API_KEY",
headers=H, json={"encrypted_value": encrypted, "key_id": pk["key_id"]}, timeout=15)
print(f"GH secret PUT: HTTP {r.status_code}")
sys.exit(0 if r.ok else 1)
PY
Recovery — if vault is stale, Heroku/GH have valid tokens
This is the inverse drift: vault rotted but the live apps still work. Less common.
- Read the live token from one of the Heroku apps:
bash heroku config:get HEROKU_API_KEY -a raxx-console-prod - Validate it (same
requests.get /accounttest as above) - Write back to vault:
bash heroku run --app raxx-console-prod --no-tty 'python -c " from app.services import vault vault.store_secret_version(\"HEROKU_API_KEY\", \"<paste here>\") "'
Prevention
Historical (#4544): the bullets below describe the pre-2026-07-10 GHA-era model (vault / GH Actions secret / Heroku config, kept in lockstep by the Mode A rotation handler). GitHub Actions deploys were retired (ADR-0135/ADR-0136); there is no GH Actions secret in this picture anymore — see "What 'drift' means here" above for the copies that actually matter today (vault / 4 Heroku config vars / session bootstrap).
- The Mode A rotation handler (#885 / PR #906 / PR #887) keeps the vault and Heroku config-var copies in lockstep on every rotation. Use it rather than manual rotation.
- Once #925 lands (
GITHUB_API_SECRETS_TOKEN), the handler can fully self-heal the (now legacy) GH-secret destination described in the "Recovery — if vault is valid, GH is stale" section above — that section is retained only for the narrow non-CI consumer case, not for Woodpecker deploys. - Audit vault / the 4 Heroku config vars / session bootstrap regularly via the validator scheduler (
console/app/services/handler_validator.py).
Postmortem template (use after every drift incident)
### HEROKU_API_KEY drift — <UTC timestamp>
- Detected: <how — log line / failed deploy>
- Vault state: <valid/stale>
- Session bootstrap state: <valid/stale>
- Heroku config state: <valid/stale on each of the 4 apps>
- Root cause: <manual rotation? failed Mode A handler? UI dashboard rotation?>
- Recovery: <Path A / Path B / inverse>
- Time to recover: <minutes>
- Followup: <issue number>
Save postmortems at docs/ops/postmortems/heroku-key-drift-<YYYY-MM-DD>.md.
Refs
- 2026-05-03 03:48 UTC incident — first observed; recovery via Path A. Cross-PR collateral: #925 filed for the durable Path B fix.
- PR #906 — Heroku Mode A handler HTTP rewrite (closes the rotation-time drift gap).
- #891 — original handler bug (CLI shell-out) that motivated #906.