fix(control,ui): close review findings - heartbeat starvation, resume race, stale UI writes

Fixes from the post-migration code review (3 major, 4 minor):

- worker: heartbeat between every drained command so a long queue can't
  starve proof-of-life into a false reap; worker_stale_after default
  60s -> 300s (one slow sync/login command must not look like a crash)
- repository.create_run: enforces one running run per agent atomically
  (agent-row FOR UPDATE on Postgres; SQLite's single writer suffices) -
  two workers claiming resumes for the same agent can no longer both
  launch claude on one session; resume surfaces the loss loudly
- headless._settle: upload the final session archive BEFORE marking the
  run finished - a resume claimed the instant a run leaves 'running'
  materializes from session_archives, and the old order let it race an
  incomplete archive into needless context re-injection (found as a
  test flake, real in production)
- store.tsx: generation token drops in-flight loadRun writes after the
  user switches runs (run A's events/log/checkmark no longer land on
  run B), plus id-keyed dedup on event appends from overlapping polls
- credsync: credential files written 0600 from the first byte
- headless: seq counter locked (reader thread + supervisor both emit
  events); proc.stdout closed after reader join
- login: submit pins to the latest CLAIMED login_start (a still-running
  one previously pinned to the wrong worker)

Suite 296 green (new: create_run conflict coverage); reaper tests track
the new staleness default.
This commit is contained in:
2026-07-21 23:41:36 -04:00
parent 1517e4dca8
commit 478d178542
18 changed files with 145 additions and 57 deletions
+14 -2
View File
@@ -55,17 +55,29 @@ def test_cancel_request_roundtrip(conn):
def test_list_running_runs_scoped_by_worker(conn):
agent = _agent(conn)
r1 = repo.create_run(conn, agent["id"], "s1", "worker-a", "spawn")
r2 = repo.create_run(conn, agent["id"], "s2", "worker-b", "spawn")
repo.finish_run(conn, r1["id"], "completed")
r2 = repo.create_run(conn, agent["id"], "s2", "worker-b", "spawn")
running = repo.list_running_runs(conn)
assert [r["id"] for r in running] == [r2["id"]]
assert repo.list_running_runs(conn, worker_id="worker-a") == []
assert [r["id"] for r in repo.list_running_runs(conn, worker_id="worker-b")] == [r2["id"]]
def test_create_run_refuses_concurrent_run_for_agent(conn):
"""One running run per agent, atomically — two workers racing a resume must not both
launch a claude process on the same session."""
import pytest
agent = _agent(conn)
repo.create_run(conn, agent["id"], "s1", "worker-a", "spawn")
with pytest.raises(repo.RunConflictError):
repo.create_run(conn, agent["id"], "s1", "worker-b", "resume")
def test_latest_run_and_agent_session(conn):
agent = _agent(conn)
repo.create_run(conn, agent["id"], "s1", "w", "spawn")
first = repo.create_run(conn, agent["id"], "s1", "w", "spawn")
repo.finish_run(conn, first["id"], "completed")
latest = repo.create_run(conn, agent["id"], "s1", "w", "resume")
assert repo.get_latest_run(conn, agent["id"])["id"] == latest["id"]