Fixes from the post-migration code review (3 major, 4 minor):
- worker: heartbeat between every drained command so a long queue can't
starve proof-of-life into a false reap; worker_stale_after default
60s -> 300s (one slow sync/login command must not look like a crash)
- repository.create_run: enforces one running run per agent atomically
(agent-row FOR UPDATE on Postgres; SQLite's single writer suffices) -
two workers claiming resumes for the same agent can no longer both
launch claude on one session; resume surfaces the loss loudly
- headless._settle: upload the final session archive BEFORE marking the
run finished - a resume claimed the instant a run leaves 'running'
materializes from session_archives, and the old order let it race an
incomplete archive into needless context re-injection (found as a
test flake, real in production)
- store.tsx: generation token drops in-flight loadRun writes after the
user switches runs (run A's events/log/checkmark no longer land on
run B), plus id-keyed dedup on event appends from overlapping polls
- credsync: credential files written 0600 from the first byte
- headless: seq counter locked (reader thread + supervisor both emit
events); proc.stdout closed after reader join
- login: submit pins to the latest CLAIMED login_start (a still-running
one previously pinned to the wrong worker)
Suite 296 green (new: create_run conflict coverage); reaper tests track
the new staleness default.