SIGN IN SIGN UP

fix(syscall): stop dropping host-process events for consumers that need them (#932)

* fix(syscall): stop dropping host-process events for consumers that need them

SyscallTracer.callback early-returned on containerID=="" before invoking
eventCallback at all, not just before reportSyscalls. The justifying comment
only reasoned about node-agent's own internal consumer
(EventHandlerFactory.ProcessEvent, which has its own empty-ContainerID drop
further downstream, confirmed at pkg/containerwatcher/v2/event_handler_factory.go),
so skipping eventCallback here was believed to be a safe no-op for that one
caller. But eventCallback is a caller-supplied callback, and other consumers of
this tracer do not have an equivalent drop -- specifically, a host/ECS agent
that watches non-containerized host processes (which have containerID=="") and
wires a callback that does not filter empty-containerID events. For that
consumer this early return silently dropped all host-process syscall events.

Move the containerID=="" check so it only skips reportSyscalls (which bypasses
the generic pipeline and must still guard against it explicitly). eventCallback
now runs for every decoded syscall regardless of containerID; node-agent's own
EventHandlerFactory.ProcessEvent still drops the empty-containerID ones itself,
so this is a no-op for node-agent's own pipeline.

Found by a human reviewer (jnathangreeg) on armosec/private-node-agent#548.

Adds TestSyscallTracerCallback covering: empty containerID still reaches
eventCallback but not reportSyscalls; non-empty containerID reaches both.

Docs-exempt: bug fix restoring event delivery to library consumers; no
documented behavior change for node-agent's own pipeline (internal consumer's
own empty-ContainerID drop is unchanged).

* fix(syscall): gate host-process event fan-out behind emitUnresolvedContainerEvents

Restoring the eventCallback loop for containerID=="" fixed the shared-library
semantics but reintroduced a steady-state cost for node-agent's own
consumers: advise_seccomp re-emits every map entry in full on every 5s poll
(no lookup-and-delete), so every host process and every stale entry for an
already-terminated container was being decoded and fanned out through the
whole event pipeline forever, for no benefit to node-agent itself.

Add an explicit emitUnresolvedContainerEvents bool to SyscallTracer/
NewSyscallTracer. tracer_factory.go's internal wiring passes false,
restoring node-agent's pre-this-PR steady-state behavior exactly; consumers
that want host-process events (e.g. the host/ECS agent) pass true. This
changes NewSyscallTracer's signature again; the private-node-agent call site
update is tracked separately.

Also fixes two test bugs found in review of the previous round:
- newSyscallEvent's t.Cleanup released the same pooled packet a second time
  after callback already released it via event.Release().
- Each subtest built a fresh synthetic datasource, which poisoned
  DatasourceEvent's package-level field-accessor cache (keyed by EventType +
  field name only, not by datasource) the moment two datasources' schemas
  ever diverged. The datasource is now built once and reused.

And rewrites the comment above the eventCallback loop to state the
containerID=="" rationale and the new flag's role instead of narrating git
history and referencing a private-repo issue number.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RkqzGWrCDfXSaCYHuZpuQa

Docs-exempt: follow-up fix to an already-open, already-reviewed PR; internal
tracer wiring/test-only change, and node-agent's own runtime behavior is
explicitly unchanged by design (that's the point of the new flag defaulting
to false at the only in-repo call site).
M
Matthias Bertschy committed
5116fe9e2a94880c197c5f243428b050a8a99e94
Parent: 1bbe089
Committed by GitHub <noreply@github.com> on 8/27/2026, 1:44:28 PM