Commit Graph
5 Commits
Author SHA1 Message Date
Neil 45f3512a33 feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support

Register DSH as a supervised Orca agent: catalog entry and detection for its
dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook
bridge, composer-ready prompt delivery, session resume, headless Source Control
AI, and title identity that no longer collides with Gemini's.

* fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini

DSH runs command hooks through its own shell executor, which drops every env var whose
name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and
ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto
scrub-safe aliases at spawn and restore them at the top of the DSH hook script.

Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH
pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini
classifier and the title status detector on DSH's whale, in the base module both copies
of that classifier read.

* test(mobile): repin the session-route closure for the DSH agent icon

* fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv

- findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a
  truncated write made install/remove delete the user's own rows. Pair each end with the
  nearest preceding start; regression test fails without the fix.
- The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status
  silently never appeared even with the remote hook installed.
- Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker.
- dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or
  'plugin' no longer marks a live agent pane non-interactive.
- Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home.
- Drop the duplicate README badge and revert an incidental doc reformat.

* refactor(dsh): share the managed-hooks reader and tighten the new modules

Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's,
with byte-identical private helpers. Both now call one readManagedHookEventsFromJson.

Also: one readTextOrAbsent instead of two spellings of the same read (dropping an
existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force)
instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse.
The patch-file transforms lose their index juggling for a predicate plus a filter.

* fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane

- applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty
  flow sequence got a block entry appended after it — invalid YAML that would leave DSH
  unable to load the user's own patch layer either. It now strips the token from an empty
  sequence (keeping a trailing comment) and returns null for a non-empty one; install
  reports that and changes nothing.
- The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass
  preserveMode.
- The relay PTY env never dropped inherited pane identity the way the local and daemon
  builders do, so a spawn that specified none could inherit the relay's own and every
  agent's hook would report against that pane.

* fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX

A refused patch file carries no managed region, so getStatus() fell through to a bare
not_installed with detail null — the actionable 'rewrite it as a block sequence' message
only ever reached the one-shot install() return. Export the predicate and check it first,
behind one shared message constant.

The owner-only mode assertion cannot hold on Windows, where chmod only toggles the
read-only attribute and mode & 0o777 reads 0o666 for any writable file.

* docs(readme): restore the DeepSeek Harness badge lost in the rebase

* test(mobile): repin the session-route closure to the measured 4221

Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three
modules above main's 4218 pin are not this change's — they arrived with the mobile work
after #22570 and were never repinned; the changelog records that split explicitly.

* fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready

Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s
timeout against an already-ready DSH composer, so a supervised worker never sees the agent
as ready.

Every existing tier reads the title, and DSH deliberately carries no title status: its rest
prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done`
is better evidence than any title anyway — it is the agent's own account of its own turn,
and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents
whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents.

* test(daemon): record the DSH transcript's true-colour I2 divergences

Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where
it reports 10 I2 divergences and failed the unlisted-transcript default of 0.

Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does
not restore to default — which is DSH's whale intro painting whole rows of true colour.
Verified as an upstream limitation rather than a regression by replaying against the
previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold.

* fix(dsh): return the new tui-idle verdict from the first-party done lane

Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather
than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it
returns READY_STRONG like the title/body lane above it. Re-verified the regression test
still fails without the lane.

* test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path

The relay builds a remote pane's env itself, so the alias mirroring there had no
test: removing the call left every suite green while remote DSH status silently
vanished. Both cases fail without it.

* docs(dsh): point the hook service at the integration reference

The reference doc had no inbound link from anywhere in the repo.
2026-09-27 22:44:18 -07:00
Neil 90801e2deb feat(agents): add first-class ZCode harness (#22464)
* feat(agents): add first-class ZCode harness

Add ZCode (Z.ai's `zcode` CLI) as a supervised Orca agent: managed lifecycle
hooks on local, SSH and Windows hosts; status, question and approval reporting;
synthetic status titles; session resume; orchestration worker launch options;
and desktop + mobile agent-picker registration.

Written against the newly open-sourced `zai-org/ZCode` (agent CLI 0.16.9), not
against a remembered screen:

- ZCode's hook runner writes a Claude-compatible stdin alias set, so it routes
  through the existing Claude-compatible vendor path while keeping its own
  identity in the sidebar.
- `PermissionRequest` fires only once the approval card is on screen and racing
  the user's answer, so it is proof the pane is blocked, not an auto-approval.
- ZCode's clarification tool is literally `AskUserQuestion` with Claude's
  questions/options shape, so Orca's question card renders it unchanged.
- ZCode's `hooks.enabled` defaults to false, which is why configured hooks were
  reported as never firing; the installer sets it.
- ZCode renames its own process to `zcode-cli`, so the expected foreground
  process cannot be the launch command or dispatch refuses the pane.
- ZCode emits no OSC title in any state and repaints its ASCII banner forever,
  so readiness comes from Orca's synthetic hook title and launch drafts wait on
  the composer box rather than on a quiet render window.

Three files crossed their max-lines limit, so each is split along a real seam:
command-line entrypoint parsing out of agent process recognition, skill
classification out of skill root discovery, and registry coverage out of the
remote hook installer tests.

Refs #10564

* fix(zcode): drop the session-option catalog and pin the orchestration contract

ZCode's CLI exposes no `--model` flag at all, and the session-option launch path
refuses to apply any option until a model id is chosen. A catalog therefore could
not deliver `--mode` per worker, and would have accepted `--model` only to drop
it silently. Take opencode's position instead: no catalog, so `worker-start
--model` is refused with a clear message and ZCode launches with the model from
its own config. `--mode` stays reachable through agent args, which is also how
the yolo default is applied.

Add a contract test covering the parts that make ZCode a usable worker:
dispatchable foreground process, stdin prompt delivery, the prompt staying out
of the launch command, and the composer-gated draft paste.

* refactor(zcode): reuse shared helpers and cut the harness down

No behaviour change; every ZCode test still passes.

- Use installer-utils' own `hookDefinitionHasManagedCommand` instead of
  re-walking a hook definition by hand, which also drops a local string reader.
- Share one `readZCodeEventMap` instead of keeping the same narrowing in both
  hook-settings and hook-config-json.
- Collapse five identical error returns into one `zcodeHookError` builder, and
  return early from the status branches instead of assigning through `let`.
- Split the event-to-status decision out of `normalizeZCodeEvent` into a pure
  `readZCodeTurn`, so the normalizer reads as decide-then-build and stops
  computing the tool name for events that never look at it.
- Take a script file name in `readManagedZCodeHookEvents` like its siblings,
  which removes a `Parameters<typeof …>` indirection at the call site.
- Drop the unused `ZCodeHookEvent` export and inline a single-use path helper.
- Correct a stale comment: ZCode's loader is a strict `JSON.parse`, so the
  in-place edit preserves key order and indentation, not comments.

* fix(zcode): address review — keep unmanaged event keys, correct comment, de-dupe README

- `removeZCodeManagedHooks` deleted any event key whose list ended up empty, so an
  unrelated `"Notification": []` the user wrote was removed as collateral whenever a
  managed hook elsewhere made the write happen. Only touch an event Orca actually
  owned something in; covered by a new regression test.
- The `isNewTurnEvent` comment claimed UserPromptSubmit was ZCode's only turn
  boundary while the expression below it also returned true for SessionStart. Say
  what the code does: SessionStart lands the idle boundary, UserPromptSubmit is the
  turn boundary (the Codex/Claude shape).
- ZCode appeared twice in the README's single agent-badge block; keep the
  local-icon entry the link checker validates and drop the favicon duplicate.

* docs(zcode): call out that the desktop bundle's CLI cannot open a session

From live testing on #22464: pointing `zcode` at the desktop app's bundled
`glm/zcode.cjs` installs Orca's hooks fine but then fails with
`Cannot find package '@zcode/tui'`, so the pane never opens a session. The
symptom reads as a broken harness when the CLI simply has no TUI. Say which
build to use and how to check before reporting a problem.

Reported-by: JWu527
2026-09-25 02:17:51 -07:00
Brennan Benson 9bb8836bb6 fix(agent-launch): wait longer for cold-boot Codex composer before dropping prompt (STA-3367) (#12853)
* fix(agent-launch): wait longer for cold-boot Codex composer before dropping prompt (STA-3367)

Continue-in-new-session pastes the handoff prompt once Codex renders its
composer glyph, gated on an 8s readiness budget. A cold/first-run Codex can
take longer than 8s to mount its composer, so the wait timed out and the
prompt was silently dropped into an empty terminal.

Marker-gated ready signals (Codex glyph, opencode show-cursor) are positive
proofs: the paste fires only when the marker actually renders, so a longer
budget can never paste prematurely — it only tolerates slow cold boots. Give
those signals a 20s budget while the markerless quiet-window signal keeps 8s.

* fix(agent-launch): share the composer-readiness budget across all three delivery owners (STA-3367)

The cold-boot fix was correct but landed as a single-path exception, and it
double-spent its own budget. Three follow-ups so the behavior is a system rule:

1. Split the PTY-spawn wait from the composer wait in pasteDraftWhenAgentReady.
   Both were handed the same budget, so a codex tab took up to 41s to report a
   dropped prompt. "Tab has a PTY" and "composer accepts input" are separate
   states: spawn keeps a fixed 8s, and the readiness budget now starts once the
   PTY exists, so a slow spawn can't shorten a cold composer's window.

2. Move the per-signal budget to draftPasteReadyBudgetMs() beside the shared
   readiness scanner. The budget is a property of the ready signal — only that
   module knows which signals are marker-gated — so all three delivery owners
   (renderer tab paste, renderer startup paste, main runtime startup paste)
   consume one policy instead of three hardcoded 8s constants.

3. Give the main-runtime startup paste the process-ownership fallback both
   renderer paths already have. It resolved null on budget expiry, silently
   dropping the prompt on worktree-create / CLI / remote-host delivery — the
   same STA-3367 failure, on the path the original fix didn't reach.

Adds coverage for the main-runtime waiter, which had none.

Test: vitest src/main/runtime src/shared src/renderer/src/lib
      src/renderer/src/components/terminal-pane — all green; tsc clean.

* test(agent-launch): consume the shared readiness budget instead of restating it

Hardcoding 20000 in the runtime waiter test meant it would keep passing if
OrcaRuntimeService stopped consuming draftPasteReadyBudgetMs — the exact drift
this PR exists to prevent. The literal values stay pinned once, in the scanner
test.

* refactor(agent-launch): collapse the readiness budget to one flat timeout

The per-signal budget (marker 20s / quiet-window 8s) tied the timeout to how
readiness is DETECTED. The budget is really a property of how slowly an agent
can boot — a marker, a quiet window, and a process check all wait out the same
cold start — so one number covers all three signals.

Replaces draftPasteReadyBudgetMs() with DRAFT_PASTE_READY_TIMEOUT_MS: drops a
constant, a branch, and two tests, and removes the only reason a delivery path
needed to know which signal class it was using.

Cost: a launch that never emits DECSET 2004 now surfaces its 'prompt not sent'
toast at 20s instead of 8s. That is the failed-launch path only; successful
markerless delivery still resolves on the 1.5s quiet window as before.

* fix(agent-launch): constrain cold Codex readiness budget

* fix(agent-launch): observe Codex readiness from PTY bind

* fix(agent-launch): anchor early Codex prompt to TUI screen
2026-08-14 13:33:05 -07:00
Neil 78f434dd85 fix(agents): deliver grok launch drafts on its composer frame (8s → 0.7s) (#13308)
* fix(agents): deliver grok launch drafts on its composer frame

Grok has no --prefill-style flag, so a launch draft (e.g. the issue URL of
a worktree created from a GitHub issue) always goes through Orca's
paste-after-ready path. That path used the default readiness signal: DECSET
2004 plus 1.5s of PTY silence. Grok shimmers its startup logo at ~12fps
until the session opens, so the quiet window never settled and the draft
fell through to the 8s hard timeout before it appeared in the composer.

Gate grok on its own composer glyph instead, anchored on the alternate-screen
switch rather than DECSET 2004: the shell that runs the launch command emits
2004 too, and its prompt may itself be the same glyph (starship, pure), so a
Codex-style anchor could paste into the shell. Grok keeps the quiet window
armed as a fallback because it renders differentially and paints the glyph
once, so a late-attaching scanner would otherwise wait out the hard timeout.

Measured against grok 1.0.0 driving the real scanner over a zsh -> grok PTY:
draft delivery moves from 8003ms to 689ms, with the URL landing unsubmitted
in the composer exactly as before.

* fix(agents): keep grok's quiet-window floor on DECSET 2004

The composer-glyph marker is anchored on the alternate-screen switch, but grok
can render inline (`--no-alt-screen`, `--minimal`, `[ui] screen_mode =
"minimal"`), where 1049h never arrives. Anchoring the quiet-window fallback
there too left those launches with no delivery path at all: readiness never
resolved, and the main-process caller drops the draft when it resolves null —
so the issue URL vanished instead of arriving late.

Give the signal two independent anchors: the marker still waits for the
alt-screen switch (so a starship/pure shell prompt can't trip it), while the
quiet window arms off DECSET 2004 exactly as the default signal does. Inline and
legacy-Windows-console launches keep their pre-existing timing; alt-screen
launches keep the fast marker path.

Verified on grok 1.0.0 over a real zsh -> grok PTY: alt-screen delivers at 687ms
via the marker, inline at 1949ms via the quiet window (the default signal
measures 1861ms on the same launch), URL landing unsubmitted in both. Adds a
recorded inline-mode trace fixture so the no-1049h path stays covered.

* fix(agents): revoke grok's alt-screen anchor when the screen is handed back

The composer-glyph anchor latched forever: once \x1b[?1049h had been seen, any
later `❯` counted as grok's composer. Two ways that pastes the launch draft into
the user's shell instead of into grok:

  - grok enters the alternate screen and then dies before painting a composer;
    the shell prompt that follows is `❯` under starship or pure.
  - a pager or editor started from the user's shell rc enters and leaves the
    alternate screen before grok is ever launched, arming the anchor against the
    shell's own prompt.

Track the anchor in stream order instead of as a latch: \x1b[?1049l revokes it,
re-entering re-arms it, and a marker only counts inside a segment where the
anchor is actually held. The chunk is walked segment by segment so ordering
within a single PTY packet is honored, with a 7-char carry — one short of the
escape sequence — so a split sequence rejoins without re-walking scanned output
into a second transition. Signals with no `markerAnchorEnd` (codex, opencode,
the default) keep their existing latch semantics untouched.

Also makes the trace-replay test model the hard timeout: the real waiters settle
at 8s, so a marker landing after that is not a delivery time.
2026-08-09 00:06:55 -07:00
Brennan Benson 973360cfce Reliably deliver the issue prompt into opencode's composer (#6798) 2026-06-30 01:16:55 -07:00