fix(containerprofilemanager): repair timestamp chain on LRU eviction and retry exhaustion (#884)
* fix(containerprofile): repair the report chain on LRU eviction and MaxAttempts exhaustion (#871) enforceMaxSize and the MaxAttempts drop path discarded a queued chunk with no replacement, forking the container's report-timestamp chain and hanging the profile in Learning forever with no field-visible symptom. Both paths now go through the same stitch-repair mechanism #866 introduced for its own split-chunk drops: MaxAttempts exhaustion routes through dropChunk, and enforceMaxSize replaces an evicted item with a stitch instead of discarding it. Since replacing every eviction with a same-count stitch makes no net progress toward MaxQueueSize by itself, enforceMaxSize bounds the in-flight stitch backlog (10% of MaxQueueSize, floor 1); once exhausted, further evictions fall back to the original unrepaired drop rather than let the queue grow without bound or convert itself entirely into stitches in one call. Signed-off-by: aryanghai12 <aryanghai1205@gmail.com> * fix(containerprofile): bound stitch admission in dropChunk and requeueSplit too enforceMaxSize checked the in-flight stitch backlog before repairing an eviction, but dropChunk and requeueSplit's lost-first-half case built and enqueued their own replacement stitch unconditionally. Left unfixed, a sustained run of MaxAttempts exhaustions or unsplittable drops could grow the backlog past the bound the eviction loop itself relies on to terminate promptly, since only enforceMaxSize's own path was gated. Factor the check into stitchBacklogFull and call it from all three stitch-admission sites; once the backlog is spent, dropChunk and requeueSplit now fall back to the same unrepaired, forked drop (dropReasonStitchBacklogExhausted) enforceMaxSize already used. Also tightens the docs page's claim that dropReasonStitchRejected is the only drop path that can leave the chain forked - dropReasonLRUBacklogExhausted already qualified that, and now dropReasonStitchBacklogExhausted does too. Signed-off-by: aryanghai12 <aryanghai1205@gmail.com> * fix(containerprofile): clamp stitch backlog, fix double-counted drops, fresh retry budget for MaxAttempts stitches Addresses review findings on the queue-drop-lineage fix: - stitchBacklog started at zero every process start but the disk queue can already hold stitches from a prior run; decrementing for one of those (which never incremented the fresh counter) drove it negative permanently, silently widening maxStitchBacklog by up to maxQueueSize. releaseStitch clamps the decrement at zero. - chunksDropped was incremented twice for a single lost chunk whenever its repair also failed (dropChunk, requeueSplit's lost-first-half case, and enforceMaxSize's failed stitch enqueue all double-counted). Each site now increments the counter once per chunk and reports both drop reasons as separate metric samples instead. - A stitch repairing a MaxAttempts-exhaustion drop inherited the parent's already-exhausted Attempts, so it got exactly one try - a no-op against a storage outage that hasn't recovered. dropChunk now gives that stitch a fresh retry budget (newStitchFor's freshAttempts); a stitch is never re-stitched, so total exposure stays bounded at two retry windows. - dropChunk, requeueSplit's lost-first-half stitch, and enforceMaxSize's own eviction loop now all gate on stitchBacklogFull before admitting a stitch, not just enforceMaxSize's loop. - NewQueueData's startup enforceMaxSize sweep now runs under qd.mu, matching the invariant enqueueStitchNoEvict's doc comment already claimed. - Documented two residual trade-offs: the MaxAttempts repair is best-effort under a sustained outage, and a restart at capacity can now shed two deltas instead of one. Signed-off-by: aryanghai12 <aryanghai1205@gmail.com> --------- Signed-off-by: aryanghai12 <aryanghai1205@gmail.com>
A
Aryan Ghai committed
b33eb186ffe5b5e09cf27682b44afcb6330bb0dc
Parent: b17e274
Committed by GitHub <noreply@github.com>
on 8/7/2026, 8:44:36 AM