Add the pi harness: lightweight local-model agents with full gate parity

Model backend rows gain a harness column (claude | pi). A pi-harness row runs
the agent through the pi coding agent instead of the claude binary — pi speaks
the OpenAI Completions API natively, so a bare vLLM/llama.cpp/Ollama endpoint
needs no LiteLLM/claude-code-router translation proxy, and the loop is far
lighter for slow local token throughput. The Claude subscription and existing
claude-harness backends are untouched.

Parity comes from generated per-agent artifacts under ~/.handler-pi (outside
the repo tree, so the clean-tree gate never trips): models.json + settings.json
render the row as a pi provider pinned as the default model; a bundled bridge
extension (pi_bridge.ts) adapts pi's events to the exact stdin/stdout contract
of `python -m handler.hooks` — the Stop/completion gate re-prompts pi with
blockers via a follow-up message, git push runs the test/build/approval gates
and denies on failure, questions defer through an ask_operator tool into the
normal answer/resume flow, and memory recall is injected at session start. The
memory tools are registered natively (pi has no MCP), shelling to a new
`python -m handler.mcpserver --call <tool>` seam that reuses the MCP server's
implementations. Skills reuse the same ~/.claude/skills sync (pi implements the
same SKILL.md standard) plus the repo's committed .claude/skills.

Sessions are single JSONL files pre-assigned via --session, so cross-worker
resume archives/materializes exactly like claude's; the prompt travels on stdin
(pi has no -- separator). The supervisor normalizes pi's event stream on the
fly: assistant message_end feeds last_output, the final agent_end becomes the
run result. The whole chain was validated live against pi 0.84.1 with a stub
OpenAI endpoint: memory injection, push-gate denial (including the protected-
branch approval gate), stop-gate block loop, and ask_operator pause all ran
end to end through the real hooks and DB.

Also: harness selector in the dashboard Models form, pi baked into the control
image (NodeSource 22 for pi's node >= 22.19 floor), PI_BIN override, docs in
docs/local-models.md, fake_pi fixture + 12 tests (361 total green).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KdGv3u3DfTsP1S188KDhVH
This commit is contained in:
Claude
2026-08-12 18:21:35 +00:00
parent 356fa276a2
commit a3a5c272a2
24 changed files with 1417 additions and 84 deletions
+24 -7
View File
@@ -506,11 +506,17 @@ const emptyModel = {
base_url: "",
model: "",
small_fast_model: "",
harness: "claude" as "claude" | "pi",
api_key: "",
env: "",
enabled: true,
};
const HARNESS_OPTS = [
{ value: "claude", label: "claude — Anthropic-compatible endpoint (LiteLLM/gateway)" },
{ value: "pi", label: "pi — bare OpenAI-compatible endpoint (vLLM/llama.cpp/Ollama)" },
];
function ModelsPanel() {
const s = useDashboard();
const [form, setForm] = useState(emptyModel);
@@ -527,6 +533,7 @@ function ModelsPanel() {
base_url: form.base_url.trim(),
model: form.model.trim(),
small_fast_model: form.small_fast_model.trim() || null,
harness: form.harness,
api_key: form.api_key.trim() || null,
env: parseKeyValues(form.env),
enabled: form.enabled,
@@ -544,6 +551,7 @@ function ModelsPanel() {
base_url: m.base_url,
model: m.model,
small_fast_model: m.small_fast_model ?? "",
harness: m.harness ?? "claude",
api_key: "", // write-only; blank = keep the stored key
env: formatKeyValues(m.env),
enabled: m.enabled,
@@ -555,13 +563,15 @@ function ModelsPanel() {
<>
<div className="faint" style={{ fontSize: "var(--text-sm)", marginBottom: 14 }}>
Alternative model backends the spawn dropdown offers next to the Claude
subscription the same <span className="mono">claude</span> binary pointed at a
different endpoint via <span className="mono">ANTHROPIC_BASE_URL</span>, so
skills, connectors, hooks, and gates apply unchanged. The endpoint must speak the{" "}
<b>Anthropic Messages API including tool use</b> a bare OpenAI-compatible
server (Ollama, llama.cpp, LM Studio) breaks tool calling; front it with LiteLLM
or claude-code-router and enable the backend&apos;s native tool parser. See{" "}
<span className="mono">docs/local-models.md</span> for working Qwen-Coder stacks.
subscription. The <b>claude</b> harness is the same{" "}
<span className="mono">claude</span> binary pointed at a different endpoint via{" "}
<span className="mono">ANTHROPIC_BASE_URL</span> the endpoint must speak the{" "}
<b>Anthropic Messages API including tool use</b>, so front a bare
OpenAI-compatible server with LiteLLM or claude-code-router. The <b>pi</b>{" "}
harness runs the lightweight pi coding agent instead, which speaks{" "}
<b>OpenAI-compatible endpoints natively</b> (vLLM, llama.cpp, Ollama no
translation proxy) hooks, gates, memory, and skills still apply via
handler&apos;s bridge. See <span className="mono">docs/local-models.md</span>.
</div>
<Card>
<div className="card-head" style={{ marginBottom: 14 }}>
@@ -594,6 +604,12 @@ function ModelsPanel() {
onChange={(v) => setForm({ ...form, small_fast_model: v })}
placeholder="qwen3-1.7b"
/>
<Select
label="Harness (which agent binary runs against this endpoint)"
value={form.harness}
onChange={(v) => setForm({ ...form, harness: v as "claude" | "pi" })}
options={HARNESS_OPTS}
/>
<Input
label={editingId != null ? "API key (blank = keep stored key)" : "API key (optional)"}
value={form.api_key}
@@ -635,6 +651,7 @@ function ModelsPanel() {
{m.name}
</span>
<div className="hstack">
{m.harness === "pi" && <Badge tone="info">pi harness</Badge>}
{m.has_api_key && <Badge tone="info">key stored</Badge>}
<Badge tone={m.enabled ? "success" : "neutral"}>
{m.enabled ? "enabled" : "disabled"}