SIGN IN SIGN UP

fix(runner): time-box record ingest and keep the incomplete-records marker working

Two small, independent robustness fixes in the session record persistence path, from investigating the 2026-09-28 EU runner OOM incident (production-session-investigation).

1. Bound each ingest POST with a per-request timeout (AbortSignal.timeout, env AGENTA_RECORDS_INGEST_TIMEOUT_MS, default 30s). Without it a single stalled request holds the per-session persist chain, and behind it the turn-end drain, open for as long as the socket stays open. A timeout now throws and is retried like any other transient error.

2. Stop flush() from consuming the per-session drop count. The turn-end finally in server.ts is the authoritative reader that marks a session incomplete when a record was dropped; flush() consumed and cleared the count first, so that reader always saw zero and the session was never marked. flush() now drains only and leaves the single read to the caller.

Unit tests updated: the flush test now asserts the count survives for the caller. Typecheck clean; runner unit suite green apart from 4 pre-existing gateway-gating failures unrelated to this change.

NOTE: the larger runaway-output memory bound (the actual OOM cause: unbounded events[] and streamed-text accumulators in tracing/otel.ts) is a separate follow-up, to be placed at the otel choke point and terminalized through the existing run-limits RUN_LIMIT_TRIPPED path, pending a recorded full-stack QA pass.
M
Mahmoud Mabrouk committed
4ff0da37da0facf26164ffeb958dbd0dc2d03fcc
Parent: 99f85e3