SIGN IN SIGN UP

feat(processtree): populate the process start time (SUB-7845) (#873)

* feat(processtree): read process start time from procfs stat field 22 (SUB-7845)

Docs-exempt: additive field population with no consumer yet; no documented
behaviour changes. Feature doc lands with the streaming attribution work
that gives the value meaning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* feat(processtree): carry procfs start time through the event pipeline (SUB-7845)

Adds StartTimeWall to events.ProcfsEvent and copies it in the ProcfsTracer and
convertProcfsEvent. convertExecEvent/convertForkEvent/convertExitEvent are
deliberately untouched: their StartTimeNs is an event wall-clock timestamp
(epoch ns, not boot ns) that only feeds pending-exit sort ordering.

Docs-exempt: additive field plumbing, no documented behaviour changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* feat(processtree): record boot-relative start time in creator side map (SUB-7845)

handleProcfsEvent stores /proc field 22's boot-relative nanoseconds in a new
pid-keyed side map and stamps the derived wall-clock value on the tree node.
The side map is the sole identity source; Process.StartTime is display-only and
inherits btime's whole-second skew.

exitByPid reclaims the side-map entry on both paths — next to processMap.Delete
and in the early return where the node is already gone.

Docs-exempt: additive, nothing reads the values yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* feat(processtree): on-demand start-time read at fork/exec node creation (SUB-7845)

Scan-only population leaves every process shorter than the 30s scan interval at
a zero start time, and short-lived processes are both the bulk of beacon-style
connections and the drivers of pid churn. ensureStartTime reads the same kernel
source as the periodic scan (/proc/<pid>/stat field 22) when a fork or exec
creates a node, so those processes get an identity too.

One read per node creation, never per event. A failed read (process already
gone) leaves zero rather than guessing. event.StartTimeNs is never consulted:
for fork/exec/exit it is the event's epoch wall-clock timestamp, a different
clock domain — pinned by TestHandleForkEvent_IgnoresEventStartTimeNs.

nsPerTick is defined per-package (feeder and creator); both are package-private
and cross-referenced in comments rather than lifted into a shared package.

Docs-exempt: additive, nothing reads the values yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* feat(processtree): expose boot-relative start time on the manager (SUB-7845)

GetProcessBootTimeNs is the method the network-stream attribution work calls to
build a per-connection process reference. The doc comment carries the identity
contract: this is the sole identity source, and the wall-clock
armotypes.Process.StartTime on tree nodes is display-only.

Docs-exempt: additive accessor, no caller yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* feat(processtree): carry StartTime through branch/copy/enrich (SUB-7845)

This is the step that makes the populated value visible downstream at all. All
three functions build a node from an explicit field list and silently strip
anything absent from it, and all three omitted StartTime:

  - buildBranchToShim  — feeds every alert branch and every stream tree
  - CopyProcess        — alert bulk manager's merged tree
  - EnrichProcess      — alert bulk manager's merge of overlapping chains

Each site gets its own named test so a future field-strip regression is caught
by name rather than showing up as a silently empty value on the wire.

Docs-exempt: additive field propagation, no documented behaviour changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* test(processtree): end-to-end guard that alert branches carry StartTime (SUB-7845)

Walks the real production path — procfs event -> ReportEvent -> creator ->
buildBranchToShim -> GetContainerProcessTree — against a creator-populated tree
rather than a hand-built fixture. Verified to fail when the branch builder stops
carrying StartTime.

Docs-exempt: test-only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* test(processtree): cover the real procfs start-time reader; harden tick pinning (SUB-7845)

Review found newProcfsStartTimeReader had zero coverage — every on-demand test
injects a fake reader — so the creator's nsPerTick could drift from the feeder's
and give one process two identities an order of magnitude apart, each internally
consistent. Verified: mutating it to 10^6 left the whole suite green.

Both conversions are now pinned against an independently read field 22 times the
contract's literal 10^7, so drift fails deterministically. The previous checks
were weaker than they looked: `ns % 10^7 == 0` only catches a 10x error when
ticks%10 != 0, and the wall-clock proximity check loses sensitivity on a
freshly booted node.

Also in the reader: resolve /proc once instead of re-stating the mount point on
every call, and log the two btime/mount failure modes that previously degraded
silently.

Docs-exempt: tests plus logging and a redundant-syscall removal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* fix(processtree): a recycled pid must not inherit the dead process's start time (SUB-7845)

A fork's pid is newborn, so a surviving side-map entry can only belong to a
process the kernel already recycled the pid away from. Exits linger for
exitCleanup.cleanupDelay (5 minutes by default), so the stale entry is readily
reachable: A exits on pid 4242, the kernel hands 4242 to B, B's fork event
arrives, and B reports A's creation time. If B lives less than one 30s scan
interval — the short-lived population the on-demand read exists for — the
periodic scan never corrects it.

That is worse than the zero it replaced: a consumer joining on (pid, startTime)
concludes A and B are the same process, which is exactly the inference this
field exists to prevent.

Scoped to the side map this change introduces, so it stays additive. This is
NOT pid-reuse hardening (SUB-7846): the shared tree node still carries the dead
process's comm, cmdline and path.

Docs-exempt: correctness fix to a field with no consumer yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* docs(processtree): document the process start time feature (SUB-7845)

Covers the boot-relative vs wall-clock split and why they are not
interchangeable, the single tick conversion and the rescaling trap, the two
population paths, zero-means-unknown and the accepted coverage gaps, the
recycled-pid guard and what it deliberately does not cover, the three copy
functions that must carry any new Process field, and the omitempty-on-a-struct
detail that makes this a changed wire value rather than a new one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* fix(processtree): restamp the display start time when a fork reuses a pid (SUB-7845)

Review catch. The recycled-pid guard cleared the boot-relative side-map entry
but left the dead process's wall-clock Process.StartTime on the reused node,
because ensureStartTime only assigns the display value when it is zero. For
exactly the case the guard exists to handle, the identity value and the value
shown to a human disagreed.

Moved the reset to where reuse is actually detected — an existing node on a
fork event — which also covers the narrower case where the side-map entry is
already gone but the node survives. The test now asserts the node's display
value, not just the accessor; it previously missed this.

Also documents that (pid, startTimeNs) is not unique: the 10ms tick
quantization leaves a residual collision when a pid is recycled inside one
tick, which matters to consumers building a join key from the tuple.

Docs-exempt: feature doc updated in the same commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* perf(processtree): take the fork path's /proc read outside the tree lock (SUB-7845)

Review catch. The design accepted this read under pt.mutex on the basis that it
is skipped when the value is already known — but the recycled-pid guard drops
the entry on every fork, so the fork path always reads and that shortcut no
longer applies to it. The cost analysed and the cost that exists had diverged.

Measured: ~7.5us per /proc/<pid>/stat open-read-parse. Fork runs to thousands
per second on a busy node, so at 1,000/s that is ~7.5ms per second of hold time
on a lock every alert type contends on, and ~22ms/s at 3,000/s.

A fork always needs the value, so reading before the lock is the same number of
reads with none of them holding it — a local reordering, not a change to lock
scope or discipline. Exec keeps its read inside ensureStartTime, where the
skip-when-known check still makes it conditional.

TestHandleForkEvent_ReadsStartTimeWithoutHoldingTreeLock pins the property with
TryRLock, so a regression fails instead of deadlocking.

Docs-exempt: feature doc updated in the same commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alon <alon@armosec.io>

* fix(processtree): gate the recycled-pid wipe on a pending exit (SUB-7845)

Maintainer review. An existing node on a fork does not prove pid reuse — some
other path may simply have created it first — so wiping unconditionally
discarded a good scan-recorded start time whenever the on-demand read then
failed, leaving both values zero. Gating on pendingExits makes the signal
precise rather than heuristic, and makes the comment true.

Also documents the limitation the same review surfaced: the dead process's
pendingExits entry survives the fork, so the delayed cleanup later runs
exitByPid on what is by then the live successor's node and the pid reverts to
unknown. Verified on this head. That is the pre-existing pid-reuse behaviour
and belongs with the reuse hardening, not here.

Docs-exempt: feature doc updated in the same commit.
Signed-off-by: Alon <alon@armosec.io>

* Update pkg/processtree/creator/processtree_creator.go

Co-authored-by: Matthias Bertschy <matthias.bertschy@gmail.com>
Signed-off-by: Alon Liwsky <40373481+AlonLiwsky@users.noreply.github.com>

---------

Signed-off-by: Alon <alon@armosec.io>
Signed-off-by: Alon Liwsky <40373481+AlonLiwsky@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Matthias Bertschy <matthias.bertschy@gmail.com>
A
Alon Liwsky committed
780bbc6ede8b968fa6ad1036b59c35c741d7936f
Parent: ce05eec
Committed by GitHub <noreply@github.com> on 8/3/2026, 8:26:02 AM