How to Refuse Installs When a Required Signal Could Not Be Evaluated

Advanced 30 minutes Platform Engineers / Security Engineers in regulated environments Security & Policy

Turn on the optional fail-closed coverage gate: declare the data sources that must be evaluable, measure in warn mode, then refuse instead of allowing unchecked. Off by default.

Overview

Every risk decision Chainsaw makes rests on data sources: a malware feed, a vulnerability database, a typosquat corpus, registry metadata, the artifact bytes. Normally, when one of those cannot be consulted — an upstream is down, a feed is stale, the request carries no artifact — Chainsaw allows the install and records that the signal did not run. That is the shipped default, on every surface, and it is deliberate: a third-party outage should not halt every build in your organisation.

For some deployments that trade is wrong. If you are regulated, air-gapped, or simply cannot accept “we allowed it because we could not finish looking”, the coverage gate inverts the trade for the sources you name.

This is off by default, and Chainsaw does not fail closed as shipped. With CHAINSAW_COVERAGE_MODE unset, nothing on this page is running and behaviour is byte-identical to a deployment without the feature. Turning it on is an explicit, per-deployment decision.

What it is not

  • Not wired into CI. chainsaw pr-scan and the Chainsaw Guard GitHub Action ignore CHAINSAW_COVERAGE_* entirely. A green PR check is not evidence that your coverage posture held.
  • Not a proof of anything on a workstation. The local guard can be uninstalled or bypassed. It gets the option for defence-in-depth and local feedback; the chokepoints an organisation can rely on are the proxy, publish, and admission.
  • Not a replacement for admission’s fail-mode. The K8s admission webhook already fails closed on database unavailability with its own mechanism. This config maps onto that mechanism rather than adding a second one.

Prerequisites

  • A Chainsaw deployment you can set environment variables on: the proxy container, the publish path, the admission webhook, or a developer workstation running the CLI guard.
  • Familiarity with monitor mode. The same “measure before you enforce” discipline applies here, and skipping it is the single most common way this feature becomes an outage.

Step 1 — Decide which sources you actually need

The gate takes data source names, not risk-signal IDs. A signal like sc.provenance_verified is a scoring rule that legitimately does not fire on most packages, so “required” cannot mean “must fire”. Availability is a property of the producer, so that is what you declare.

There are eight names in configuration version 1:

SourceBacks
malwareKnown-malicious package feed (OSV / GHSA malicious-packages, Docker malware feed)
cveVulnerability matching (CVE / CVSS / EPSS, and the KEV cross-reference derived from it)
typosquatTyposquat detection
provenanceSignature / attestation verification
registry_metadataRegistry packument fetch — package age, deprecation, yank and withdrawal reads
checksumChecksum enforcement
install_scriptsInstall-script inspection
hidden_unicodeHidden-Unicode scanning

An unrecognised name is a hard configuration error, never a silent no-op. A typo fails at startup instead of leaving you with a control you believe in and do not have.

Seventeen of the twenty-five marketed signals have no source name. Publisher-diff / maintainer takeover, publish velocity, weekly downloads, maintenance status, repo liveness, code smell, capability grading, AI artifact scanning and the rest are derived or post-merge signals whose availability depends on several inputs at once. A truthful “was this evaluable?” answer for them is a dependency graph, not a lookup, so version 1 deliberately excludes them rather than reporting a status it cannot compute. If the signal you care about is not in the table above, the coverage gate is not the control for it today.

Step 2 — Check what your surface can honestly attest

A surface can only refuse on a source it actually consults. Requiring one it cannot observe would hand you a permanently inert control, so Chainsaw treats it as a startup configuration error instead.

Registry proxy and publish — metadata-only

Both evaluate the package coordinate without reading artifact bytes: the proxy scans with no artifact to keep the install hot path fast, and the publish pre-run caps at the metadata tier. They accept only:

malware, cve, typosquat, provenance, registry_metadata

Naming checksum, install_scripts or hidden_unicode there makes the server refuse to start, quoting the setting. That is intentional — those providers would report “needs artifact”, which classifies as not applicable, which never blocks, so the control would look configured and never fire.

Workstation guard — all eight, with honest refusals

The guard reads bytes when it has them, so it accepts every name. What it can truthfully report:

SourceAvailable on the guard?
typosquatAlways — the corpus is embedded and works offline.
malwareOnly after chainsaw guard update has fetched the full OpenSSF feed. Without it the guard has the embedded famous-attack floor, which is partial coverage.
cveNever offline — needs the server.
registry_metadataNever offline — needs the server.
provenanceNever offline — needs the server.
checksumOnly with staged artifacts or deep mode.
install_scriptsOnly with staged artifacts or deep mode.
hidden_unicodeOnly with staged artifacts or deep mode.

Requiring cve on an offline guard refuses every install. That is the correct answer, not a bug. The only truthful thing an offline guard can say about a vulnerability database it cannot reach is “I could not look, so I will not vouch.” If you want those sources enforced, require them on the proxy, where they can actually be evaluated.

K8s admission — a config surface, not a second gate

The webhook reads the same variables and maps them onto its existing fail-mode: closed sets fail-mode closed; off and unset contribute nothing; warn is accepted, logged, and changes nothing. required, grace and max_ledger_age are parsed and validated there — so one fleet-wide config is uniformly legal everywhere — but not consumed, because admission has no coverage ledger.

The webhook combines this with CHAINSAW_ADMISSION_FAIL_MODE by strictness: neither knob may loosen the other, so a stale CHAINSAW_ADMISSION_FAIL_MODE=open in a Helm values file cannot silently defeat a declared fail-closed posture. Every override is logged at startup.

Step 3 — Measure in warn mode first

export CHAINSAW_COVERAGE_MODE=warn
export CHAINSAW_COVERAGE_REQUIRED=malware,typosquat

warn evaluates the gate and records exactly what closed would have refused, under its own counter, and refuses nothing. This is a first-class mode, not a debug flag — an opt-in control that halts every build in your organisation during someone else’s outage gets switched off in week two, so measure your real would-block rate against production traffic before you flip.

Leave it running long enough to cover an upstream wobble. Compare the would-block count against your install volume. If the number is not a number you would accept as refusals, fix the source or shorten the required list before continuing.

Step 4 — Tune the two staleness bounds

Coverage is recorded at scan time and replayed on cache hits, so it can be wrong in both directions. Two bounds fix that, and they pull in opposite directions on purpose.

VariableDefaultWhat it does
CHAINSAW_COVERAGE_MAX_LEDGER_AGE15mThe outer bound, always wins. Coverage older than this counts as unavailable, so a scan cached while everything was healthy cannot vouch for a source during an outage.
CHAINSAW_COVERAGE_GRACE30sRescues a source that is down now but was healthy within the window, so a three-second registry hiccup does not refuse a build.

Validation rejects grace >= max_ledger_age, which would make the grace window unreachable.

grace is effectively inert on the proxy and publish surfaces. Grace can only rescue a source that has a strictly earlier healthy observation, and a single intelligence report carries one collection timestamp — the server-side producer keeps no cross-scan history. So on a server, a source’s fresh status stands on its own and grace has nothing to draw on. It is a working rail on the workstation guard. Do not plan a rollout assuming grace will absorb upstream blips on the proxy; use warn mode to size that instead.

A required source with no entry at all counts as unavailable. “Did not run” is not the same claim as “ran and succeeded”, and for a gate you asked to fail closed, the conservative reading is the only safe one.

Step 5 — Flip to closed, one surface at a time

export CHAINSAW_COVERAGE_MODE=closed
export CHAINSAW_COVERAGE_REQUIRED=malware,typosquat

What a refusal looks like:

  • Proxy / publish — HTTP 403 with error code CHW-2307, the standard ecosystem block envelope, and an install.blocked webhook carrying block_source: coverage_gate. The reason field names the sources that were missing.
  • Workstation guard — every package in the invocation is refused with severity coverage and the standard blocked exit code.

Two properties worth knowing:

  • The gate runs before, and independently of, policy evaluation, and a gate block is not overridable by a policy allow. Policy can only add blocks. That ordering is what keeps the guarantee alive when the Rego bundle is broken, absent, or stale — every policy-engine error path in Chainsaw is specified to fail open, so a fail-closed guarantee cannot live inside it.
  • Misconfiguration is fatal at startup. An explicitly configured posture that cannot be honoured refuses to boot rather than silently downgrading to off. Silently downgrading would reproduce the exact failure the option exists to prevent.

Step 6 — Keep the escape hatch reachable

CHAINSAW_COVERAGE_BREAK_GLASS=1 chainsaw npm install some-package

Break-glass disables the gate for that invocation or process and logs loudly every time. Its existence is deliberate: a security control with no documented way out gets defeated by an undocumented one. Alert on the log line rather than removing the variable.

Verification

  • With CHAINSAW_COVERAGE_MODE unset, install a package that you know has a degraded signal. It should install, exactly as before.
  • Set CHAINSAW_COVERAGE_MODE=warn with a source your surface cannot attest (for example cve on an offline guard). Confirm you get a warning and the install still proceeds.
  • Set CHAINSAW_COVERAGE_MODE=closed with the same value. Confirm the refusal, and that the reason names the source.
  • On a server, set CHAINSAW_COVERAGE_REQUIRED=checksum. Confirm the process refuses to start and the error names CHAINSAW_COVERAGE_REQUIRED.

Troubleshooting

The proxy will not start after I set the variables. Read the error — it names the setting and the offending value. The two common causes are an unknown source name and an artifact-bound source (checksum, install_scripts, hidden_unicode) required on the metadata-only server surfaces.

Everything is refused on my workstation guard. You almost certainly required a source the offline guard cannot see. cve, registry_metadata and provenance need the server; malware needs chainsaw guard update; the artifact-bound three need staged artifacts or deep mode.

A brief upstream outage refused a batch of installs on the proxy. Grace does not help there (see Step 4). Either drop the affected source from the required list, or run that surface in warn and enforce on a surface with a tighter dependency on the source.

My PR check passed but the proxy refused the same package. Expected — chainsaw pr-scan does not read CHAINSAW_COVERAGE_*.