Gitleaks runbook
System: gitleaks nightly security scan + ci-boundary post-merge scan
Owner: sre-agent
Last incident: 2026-08-10 (ci-boundary gitleaks step scanned only 1 commit on release push despite depth: 0 — see #4463/#4464, RCA docs/incidents/2026-08-12-ci-boundary-gitleaks-scan-scope.md)
Last reviewed: 2026-08-12 (#4483 — added verify-full-history regression gate to ci-boundary.yaml and security-scan-nightly.yaml, following #4481's class sweep of the same depth: 0 without partial: false gap across the rest of .woodpecker/)
How to tell it's broken
- GitHub issue filed by
security_file_issues.pywith labeltype:securityand title matchinggitleaks:— check if the flagged value is a real secret or a false positive. - Nightly scan failing with findings that were previously suppressed (allowlist regression).
- CI gitleaks step exits non-zero on a PR that should be clean.
ci-boundary(release/main push) gitleaks step log shows a suspiciously lowN commits scannedcount relative to the pushed range — see Failure mode E.
Where to find findings for a ci-boundary run
Since #4464, the gitleaks-report-upload step publishes the itemized JSON
report (secret values redacted via --redact; file/rule/commit metadata
intact) to s3://raxx-ci-artifacts/gitleaks-reports/<CI_PIPELINE_NUMBER>/gitleaks.json
on every ci-boundary run — pass or fail. Retention is 14 days. Pull it with:
aws s3 cp s3://raxx-ci-artifacts/gitleaks-reports/<pipeline-number>/gitleaks.json - \
| python3 -m json.tool
The WP step log's "leaks found: N" summary line alone is not enough to triage a finding — always pull the artifact for file/rule/commit detail rather than re-running the scan locally.
How to diagnose (in order)
- Read the issue title — it includes the file path and line number. Read that line in the repo.
- Classify the finding: is the value a real credential (API token, private key, password) or a public identifier / placeholder?
- Run gitleaks locally on the full history to reproduce:
gitleaks detect --source=. --config=.gitleaks.toml --no-git=false --report-path=/tmp/gl.json --report-format=json python3 -c "import json; [print(f['RuleID'], f['File'], f['StartLine'], repr(f['Match'])) for f in json.load(open('/tmp/gl.json'))]" - If the finding is a false positive and an allowlist entry exists, check why it isn't suppressing:
- Is
regexTarget = "match"set? Without it, regexes test against the extracted Secret (bare capture group), not the full Match string. - Does the regex actually match the Match string? Test with Python:re.search(pattern, match_string). - Is the path allowlist entry correct? Check the path prefix in the finding'sFilefield.
Known failure modes
Failure mode A: allowlist regex not suppressing — missing regexTarget = "match"
Symptom: An allowlist regex entry exists for the pattern, but gitleaks still flags it in full-history scans. The regex is written to match a variable-assignment expression (e.g. cf_access_account_id\s*=\s*"..."), not just the bare secret value.
Cause: gitleaks [allowlist] regexes default to matching against the extracted Secret field (just the capture group from the rule's regex). Without regexTarget = "match", a regex like varname\s*=\s*"value" tests against value — which doesn't match the full expression.
Fix:
1. Add regexTarget = "match" to the [allowlist] block in .gitleaks.toml. This applies to all entries in the block.
2. Verify all other regex entries still work (they should — substring match against the longer Match string is backward-compatible for most patterns).
3. Run gitleaks detect --source=. --config=.gitleaks.toml --no-git=false and confirm findings drop to 0.
Verification: Local full-history scan exits 0. File PR, confirm CI scan passes.
Failure mode B: new false positive — public identifier flagged as generic-api-key
Symptom: A hex string, UUID, or high-entropy public identifier committed to a .tf, .json, or similar config file is flagged under generic-api-key or generic-api-key.
Cause: gitleaks generic-api-key rule fires on Shannon entropy thresholds, not on pattern context. Public identifiers (Cloudflare account IDs, zone IDs, GCP project IDs, etc.) often have enough entropy to trigger it.
Decision criteria — add to allowlist only if ALL of these are true:
- The value is publicly visible (e.g. appears in dashboard URLs, API response metadata, or documentation)
- It cannot be used to authenticate — possession of the value alone grants no access
- The variable name documents the public nature (e.g. _account_id, _zone_id, not _token, _key, _secret)
Fix:
1. Add a regex entry with regexTarget = "match" (already set if the block has it) matching variable_name\s*=\s*"value_pattern".
2. Add a comment citing the issue number and explaining why it's safe.
3. Run full-history scan to verify.
4. Do NOT use a bare hex pattern as the regex — it will over-suppress any 32-char hex value.
Verification: gitleaks detect --source=. --config=.gitleaks.toml --no-git=false exits 0.
Failure mode C: real secret committed — must rotate
Symptom: gitleaks flags a value that is a live credential (API token, private key, password). The variable name or context confirms it (_token, _key, _secret, password, BEGIN PRIVATE KEY, etc.).
Fix: This is a security incident, not a false positive. Follow the security incident response protocol:
1. Classify severity (SEV-1 if token is for a production system with blast radius, SEV-2 if staging/internal).
2. Rotate the credential immediately at the vendor dashboard or API. Do not wait for history rewrite.
3. Update vault at the canonical path (/MooseQuest/<vendor>/).
4. Verify all consumers of the rotated key still work.
5. File a rewrite plan (BFG or git filter-repo) in the incident issue. Do NOT execute the rewrite without explicit operator authorization (force-push to main is destructive).
6. Write an RCA.
Verification: Vendor confirms old token is revoked. New token passes smoke test. Vault has new value.
Failure mode D: path allowlist not suppressing — wrong path prefix
Symptom: A file at console/app/sops/rotation/foo.md is flagged, but the allowlist only covers ^docs/ops/runbooks/rotation/.
Cause: The same file content lives at multiple paths (e.g. synced copies). Each distinct path prefix needs its own entry.
Fix: Add the missing path regex to the paths array in .gitleaks.toml. Keep entries narrow — prefer directory-level paths over single-file entries.
Verification: Full-history scan exits 0.
Failure mode E: ci-boundary scans anomalously few commits despite depth: 0
Symptom: WP ci-boundary gitleaks step log shows INF N commits scanned
where N is small (often 1) on a push-triggered release/main build, even
though .woodpecker/ci-boundary.yaml's clone.settings.depth is 0 (full
history, per the file's own comment).
Cause: WP's woodpeckerci/plugin-git defaults partial: true on every
event that is not a tag event. Per plugin-git's own docs (docs.md, partial
row): partial clone "only fetch[es] the one commit and it's blob objects...
overwrite[s] depth with 1." depth's own doc row confirms it is "overwritten
by partial." So a push-triggered clone silently ignores depth: 0 unless
partial: false is also set — depth: 0 alone is not sufficient.
Confirmed against woodpeckerci/plugin-git main-branch docs 2026-08-12; see
docs/incidents/2026-08-12-ci-boundary-gitleaks-scan-scope.md.
Fix: Add partial: false alongside depth: 0 in the clone.settings
block:
clone:
- name: clone
image: woodpeckerci/plugin-git
settings:
depth: 0
partial: false
This is already applied in .woodpecker/ci-boundary.yaml as of #4463. If a
new WP pipeline is added with a full-history clone requirement (any
depth: 0 push/pull_request/cron/manual-triggered pipeline — NOT a tag
event), it needs the same partial: false pairing.
Class-sweep status (#4481, closed 2026-08-12): a repo-wide grep of
.woodpecker/*.yaml for clone.settings.depth: 0 found exactly three
matches. All three now carry partial: false:
| File | partial: false added |
Behavior change? |
|---|---|---|
ci-boundary.yaml |
#4463 (2026-08-12) | Yes — root cause of this failure mode; confirmed fix intent, live-run confirmation is #4463 action item 2 (due 2026-08-26). |
security-scan-nightly.yaml |
#4481 (2026-08-12) | Yes. Trigger is cron/manual (both non-tag, both subject to partial: true's default) and the gitleaks step's --no-git=false full-history walk genuinely depends on clone depth. Before this fix the nightly gitleaks step was almost certainly scanning ~1 commit on every run, same class as the ci-boundary incident. Expect the reported commit count and finding count to jump substantially on the next scheduled (08:07 UTC) or manual run — consistent with the ~16 pre-existing matches already found via full-history local runs in #4445/#4463 (triaged separately under #4465). |
migration-collision-check.yaml |
#4481 (2026-08-12) | No observed change expected. Its check step resolves --base-sha via gh api (not git fetch) and only runs two-endpoint git diff <base-sha> HEAD / git show <base-sha>:<path> — that needs the two named commits present as objects, not a full ancestor walk. This REQUIRED status check has been passing on real migration PRs (#4294, #4297, #4322, #4369) under the pre-existing partial: true-default clone, which is direct evidence the check's logic does not depend on clone depth. partial: false was added anyway for config honesty (the file's own header comment already claimed "full history... needed") and as defense-in-depth against a future regression to a git fetch-based base-ref resolution. |
A repo-wide grep for git clone/git fetch --depth outside clone.settings
(ci-pr.yaml, queue-docker-smoke.yaml's deliberate --depth=1 shallow
fetches; deploy-prod.yaml/deploy-staging.yaml/cut-release-candidate.yaml/
review-app-console.yaml's git fetch --unshallow pattern;
vcpkg-manifest-check.yaml's in-step clone of an external repo, not this
one) confirmed none of those are instances of this class — they either
intend a shallow fetch, already use the --unshallow workaround from an
earlier incident, or aren't cloning this repo at all.
Verification: Trigger a real push-triggered ci-boundary build (next
release/main push) and confirm the gitleaks step log's N commits
scanned matches git rev-list --count <ref> for the pushed ref, not 1.
Local reproduction (does not require a live WP run, useful for pre-merge
confidence): gitleaks detect --source=. --config=.gitleaks.toml --no-git=false
against a full (non-shallow) local clone should report commit counts in the
thousands for this repo, not single digits. For security-scan-nightly.yaml,
the same local reproduction applies; live confirmation requires either
waiting for the next 08:07 UTC cron run or a manual trigger via the WP API
(POST /api/repos/.../pipelines) and reading the gitleaks step log — not
yet performed as of #4481's merge (config fix landed; live-run confirmation
is a natural follow-up on the next scheduled run, same deferral pattern as
4463's action item 2).
Automated regression gate (#4483): Both ci-boundary.yaml and
security-scan-nightly.yaml now carry a verify-full-history step,
inserted immediately after clone and before the step that consumes full
history (gitleaks). It is a one-line boolean assertion:
- name: verify-full-history
image: alpine/git:2.45.2
commands:
- |
if [ "$(git rev-parse --is-shallow-repository)" = "true" ]; then
echo "FATAL: clone is shallow despite depth:0 — plugin-git 'partial' default likely overrode depth. See docs/ops/runbooks/gitleaks.md Failure mode E." >&2
exit 1
fi
echo "Full-history clone confirmed (not shallow)."
In security-scan-nightly.yaml the step also carries failure: ignore,
matching its sibling scanner steps in that file during the A/B dual-run
pilot (ADR-0134) — the step still exits non-zero and fails loud in the WP
step log, it just doesn't escalate beyond what the rest of that pipeline
already does. ci-boundary.yaml's copy has no failure: ignore; it is a
genuine hard gate on that pipeline.
Why a boolean assertion, not a commit-count comparison. The #4463 RCA's
action item #4 originally floated comparing gitleaks' own "N commits
scanned" against git rev-list --count HEAD. That comparison is a null
check: both numbers are computed inside the same clone, so when
partial: true silently truncates the clone to 1 commit, git rev-list
--count HEAD inside that same shallow clone also reports 1 — the two sides
agree, and the check would not have caught #4463. git rev-parse
--is-shallow-repository instead reads .git/shallow's presence directly,
which is set by the clone itself regardless of how many total commits the
repo has. This eliminates the false-positive risk on a genuinely small
repo/branch history (fresh branch, single commit) without any tunable
threshold. Ruling: PM+QA+architect adjudication on #4483, 2026-08-12.
Locally-verified edge cases (required before merge, not just asserted):
- A full (non-shallow) clone of a fresh 1-commit repo: git rev-parse
--is-shallow-repository → false → step passes.
- A --depth 1 clone of the same repo (git clone --depth 1 file://...,
note plain local-path clones ignore --depth — use a file:// URL to
reproduce): git rev-parse --is-shallow-repository → true → assertion
trips, step exits 1.
migration-collision-check.yaml intentionally does not get this step
(or any shallow-clone check). Its check step resolves the PR base SHA via
gh api, not git fetch, and its git diff/git show calls only need the
two named commits present as objects — not a full ancestor walk. A
shallowness assertion there would test a precondition the pipeline's logic
never consumes. See the migration-collision-check.yaml row in the class-
sweep table above for the evidence trail (#4294, #4297, #4322, #4369 passed
under the pre-existing partial: true default).
Testing an allowlist change
Always test allowlist changes against the full-history scan, not just --no-git:
gitleaks detect --source=. --config=.gitleaks.toml --no-git=false
# Expected: "no leaks found" (exit 0)
For CI v8.18.4 (the pinned version in nightly-security-scan.yml), verify with:
curl -sSL https://github.com/gitleaks/gitleaks/releases/download/v8.18.4/gitleaks_8.18.4_darwin_arm64.tar.gz \
| tar -xz -C /tmp gitleaks
/tmp/gitleaks detect --source=. --config=.gitleaks.toml --no-git=false
Emergency stop
To disable the nightly scan temporarily (e.g. while a real incident is being triaged):
- Edit
.github/workflows/nightly-security-scan.yml— comment out theschedule:trigger. - Open a PR with the change. Do not merge to main without a paired "re-enable" commit ready.
- Log the disable in the incident issue with a stated re-enable time.
Escalation
Escalate to Kristerpher when:
- A confirmed real secret is found in history and rotation has been completed but force-push authorization is needed for history rewrite
- The CI scanner version (v8.18.4) needs to be bumped — check gitleaks CHANGELOG for breaking changes to allowlist behavior first
- A new finding cannot be classified as real vs false positive within 30 minutes of investigation
- The gitleaks-report-upload step in ci-boundary.yaml is failing to
authenticate to AWS: the raxx-ci-boundary-gitleaks-report-writer-ci IAM
user (infra/ci/ci-artifacts.tf, #4464) requires a live terraform apply
(with real terraform.tfvars, per infra/ci/main.tf's documented Apply
order) plus scripts/ops/secrets/mint_ci_aws_keys.sh to mint access keys
into vault at /MooseQuest/aws/ci-boundary/gitleaks-report-writer/ before
this step can succeed for the first time — this is a deliberate,
operator-gated step per the credentials-lockdown posture (new AWS key
mints escalate), not something sre-agent applies unattended.