fix(syscall): stop dropping host-process events for consumers that need them (#932)
* fix(syscall): stop dropping host-process events for consumers that need them SyscallTracer.callback early-returned on containerID=="" before invoking eventCallback at all, not just before reportSyscalls. The justifying comment only reasoned about node-agent's own internal consumer (EventHandlerFactory.ProcessEvent, which has its own empty-ContainerID drop further downstream, confirmed at pkg/containerwatcher/v2/event_handler_factory.go), so skipping eventCallback here was believed to be a safe no-op for that one caller. But eventCallback is a caller-supplied callback, and other consumers of this tracer do not have an equivalent drop -- specifically, a host/ECS agent that watches non-containerized host processes (which have containerID=="") and wires a callback that does not filter empty-containerID events. For that consumer this early return silently dropped all host-process syscall events. Move the containerID=="" check so it only skips reportSyscalls (which bypasses the generic pipeline and must still guard against it explicitly). eventCallback now runs for every decoded syscall regardless of containerID; node-agent's own EventHandlerFactory.ProcessEvent still drops the empty-containerID ones itself, so this is a no-op for node-agent's own pipeline. Found by a human reviewer (jnathangreeg) on armosec/private-node-agent#548. Adds TestSyscallTracerCallback covering: empty containerID still reaches eventCallback but not reportSyscalls; non-empty containerID reaches both. Docs-exempt: bug fix restoring event delivery to library consumers; no documented behavior change for node-agent's own pipeline (internal consumer's own empty-ContainerID drop is unchanged). * fix(syscall): gate host-process event fan-out behind emitUnresolvedContainerEvents Restoring the eventCallback loop for containerID=="" fixed the shared-library semantics but reintroduced a steady-state cost for node-agent's own consumers: advise_seccomp re-emits every map entry in full on every 5s poll (no lookup-and-delete), so every host process and every stale entry for an already-terminated container was being decoded and fanned out through the whole event pipeline forever, for no benefit to node-agent itself. Add an explicit emitUnresolvedContainerEvents bool to SyscallTracer/ NewSyscallTracer. tracer_factory.go's internal wiring passes false, restoring node-agent's pre-this-PR steady-state behavior exactly; consumers that want host-process events (e.g. the host/ECS agent) pass true. This changes NewSyscallTracer's signature again; the private-node-agent call site update is tracked separately. Also fixes two test bugs found in review of the previous round: - newSyscallEvent's t.Cleanup released the same pooled packet a second time after callback already released it via event.Release(). - Each subtest built a fresh synthetic datasource, which poisoned DatasourceEvent's package-level field-accessor cache (keyed by EventType + field name only, not by datasource) the moment two datasources' schemas ever diverged. The datasource is now built once and reused. And rewrites the comment above the eventCallback loop to state the containerID=="" rationale and the new flag's role instead of narrating git history and referencing a private-repo issue number. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RkqzGWrCDfXSaCYHuZpuQa Docs-exempt: follow-up fix to an already-open, already-reviewed PR; internal tracer wiring/test-only change, and node-agent's own runtime behavior is explicitly unchanged by design (that's the point of the new flag defaulting to false at the only in-repo call site).
M
Matthias Bertschy committed
5116fe9e2a94880c197c5f243428b050a8a99e94
Parent: 1bbe089
Committed by GitHub <noreply@github.com>
on 8/27/2026, 1:44:28 PM