SIGN IN SIGN UP

fix(sbom): cap host SBOM scan parallelism to the container's CPU limit (#978)

* fix: cap host SBOM scan parallelism to the container's CPU limit

Live-cluster verification found that with sbomGenerationEnabled and
hostMonitoringEnabled both on, node-agent's own liveness probe
(GET :7888/livez) timed out ~5.5 minutes into every startup and kubelet
killed the container, repeatedly, so the host SBOMSyft CR never advanced
past status=initializing. Disabling sbomGenerationEnabled alone made the
pod stable immediately; raising the container's limits 8x on CPU
(394m -> 3) and 4x on memory did not help at all.

That last detail identifies the mechanism. Syft's Parallelism config
defaults to 0, meaning runtime.NumCPU(), and hostSbomConfig never
overrode it -- so on an 8-CPU node the host root-filesystem walk spread
cataloging across 8 goroutines while the cgroup quota was a fraction of
one CPU, and CFS throttling starved the goroutine serving /livez.
Raising the quota without capping parallelism leaves that ratio roughly
unchanged, which is why the extra resources bought nothing.

hostSbomConfig now takes a parallelism parameter and calls Syft's own
.WithParallelism(n), scoped to the host scan only -- deliberately not a
process-wide GOMAXPROCS change, which would also reshape the eBPF
tracers and event processing. n = max(1, CPU_LIMIT_MILLIS / 1000), where
CPU_LIMIT_MILLIS is supplied by the chart through the Kubernetes
downward API (resourceFieldRef: {resource: limits.cpu, divisor: "1m"}),
so 394m resolves to 1.

The downward API is used in preference to reading /sys/fs/cgroup/cpu.max
because node-agent bind-mounts the host's /sys/fs/cgroup over its own
(see pkg/metricsmanager/otel/resource_metrics.go, which had to build
container-scope cgroup resolution for exactly this reason): a root-level
cgroup read inside this container returns the node's quota, not the
container's, and its failure direction is invisible -- it falls back to
runtime.NumCPU(), i.e. exactly the unbounded behaviour being fixed.

When CPU_LIMIT_MILLIS is absent (an older chart), the scan falls back to
runtime.NumCPU() and logs that at WARN whenever host SBOM scanning is
enabled, since it means the cap is not in effect. The resolved n is
logged at Info at scan start, so a silent regression back to n == NumCPU
cannot pass an otherwise all-green live-cluster run. The new
hostSbomScanParallelism config key (default 0 = computed) forces a value
without a new image build.

The chart change lives in kubescape/helm-charts and must ship first.
The sbom-scanner sidecar has the identical exposure and is deliberately
left uncapped here: this change's live verification exercises only the
in-process path, so capping the sidecar too would ship unverified.

Full design, measurement protocol and live-cluster acceptance criteria:
.omc/plans/host-sbom-sidecar-offload-plan.md (PR 1).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Matthias Bertschy <matthias.bertschy@gmail.com>

* fix: correct Syft default-parallelism claim (NumCPU()*4, not NumCPU())

Independent code review of the parallelism-cap PR found the commit's
central mechanism claim was factually wrong for the syft actually vendored
here: go.mod replaces github.com/anchore/syft with github.com/kubescape/syft
v1.32.0-ks.2, and that fork resolves parallelism 0 to runtime.NumCPU() * 4,
not runtime.NumCPU(). All docs, comments, and the WARN fallback log message
understated the pre-fix exposure by 4x (32 goroutines on the 8-CPU test
node, not 8) and mischaracterized the NumCPU() fallback as "unbounded" when
it is itself already a 4x reduction from Syft's own default.

No runtime behavior change -- the fix (n = max(1, CPU_LIMIT_MILLIS/1000),
falling back to runtime.NumCPU()) is unaffected either way. This corrects
the explanatory narrative and the operator-facing WARN string to be
accurate about the mechanism being fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Matthias Bertschy <matthias.bertschy@gmail.com>

* docs: record live-cluster confirmation of the parallelism cap fix

Deployed quay.io/matthiasb_1/node-agent:host-pr1 to
do-fra1-matthias-host-test with sbomGenerationEnabled/hostMonitoringEnabled
on and node-agent at its original 394m/682Mi limits (not the earlier
experimentally-raised 3 CPU/3Gi). Confirmed: parallelism:1 logged at scan
start, node-agent ran 20+ minutes with zero restarts through the full scan
cycle (every prior attempt died by ~5.5 minutes), and the host SBOMSyft CR
advanced from its permanently-stuck resourceVersion=1/Initializing to
resourceVersion=2/Learning with 461,801 bytes of real content.

Full measurement recorded in the working plan document.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Matthias Bertschy <matthias.bertschy@gmail.com>

* fix: keep host SBOM fallback serial without CPU limit metadata

Signed-off-by: Matthias Bertschy <matthias.bertschy@gmail.com>

---------

Signed-off-by: Matthias Bertschy <matthias.bertschy@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
M
Matthias Bertschy committed
91f7672edc843286e5a466a508a94dcbdf4bcb4e
Parent: 9b41460
Committed by GitHub <noreply@github.com> on 9/22/2026, 11:56:15 AM