mirror of
https://github.com/stablyai/orca.git
synced 2026-09-23 00:02:29 +00:00
a3edabcd7babc3a488a35be91a30da3dcd1eeaec
255
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
53bd956ef6 |
ci(e2e): keep the relay version markers in the e2e build artifact (#15303)
build-relay.mjs writes each relay's marker as out/relay/<platform>/.version, and upload-artifact excludes dotfiles unless include-hidden-files is set. The markers were therefore stripped from e2e-build-out, so every consumer that actually starts a relay failed with: Orca's local relay build is missing its version marker at out/relay/linux-x64/.version This stayed latent because the sharded e2e lane never sets ORCA_E2E_SSH_DOCKER, so its SSH specs skip instead of touching the relay. The changed-specs lane does set it whenever a changed spec needs Docker SSH, which is why the failure only appears on PRs that touch SSH-adjacent code. |
||
|
|
66a5e5d245 |
fix(shell): repair worktree HISTFILE in plain zsh panes via one positive feature channel (#15258)
* fix(shell): repair worktree HISTFILE in plain zsh panes via one positive feature channel
A plain zsh pane — no startup command, no agent overlay — was never wrapped, so
Orca's HISTFILE repair never ran in it. macOS `/etc/zshrc` assigns
`HISTFILE=${ZDOTDIR:-$HOME}/.zsh_history` with no check-before-set and runs
before any file Orca controls, so per-worktree history was a silent no-op for
every ordinary pane on the primary platform.
Wrapping those panes needs a way to say which wrapper features a shell should
turn on. That channel is one exported variable, ORCA_SHELL_FEATURES, carrying a
comma-separated positive allowlist from a closed set (history, markers, ready,
identity, overlay). The wrapper .zshenv reads it into a plain, non-exported
array and unsets it in its first executable lines, before the user's own
.zshenv — so the selection survives .zshenv -> .zprofile -> .zshrc -> .zlogin in
this process but physically cannot reach a child. There is no negative or
suppression variable anywhere; an absent or inherited value can only ever mean
fewer features. ORCA_HISTFILE is consumed and destroyed the same way, which
removes the root cause of #11146 instead of patching it.
All order-sensitive wrapper work now lives in one `__orca_shell_epilogue`
defined in .zshenv and invoked exactly once, from .zshrc for a non-login shell
and .zlogin for a login shell, with each feature an independent guard.
Selection is a pure function of spawn env and launch intent, so a pane wrapped
only for history gets no OSC 133 and is observably identical to the unwrapped
pane it used to be.
Generation is now fail-closed: wrapper files are written to a temp name and
renamed, every required path is verified non-empty, and ZDOTDIR is only set when
that holds. Previously a failed write still pointed ZDOTDIR at an empty dir and
the user silently lost their entire zsh config.
Orca also recognised only its own `*/shell-ready/zsh` dir shape when deciding
what the user's ZDOTDIR was, so being launched from any other terminal that had
hijacked ZDOTDIR captured that as the user's config dir. Ownership is now
established positively — a stamped marker file, or Orca's own dir shape for
wrappers written by older builds — and an inherited ZDOTDIR holding no zsh
startup file is ignored. No vendor is detected by name.
* fix(shell): make the zsh epilogue option-proof and stop history widening relay wrapping
Review follow-ups on the feature-channel PR.
- `emulate -L zsh` as the epilogue's first statement. It runs after the user's
own config, so `setopt no_unset` made the precmd_functions append a fatal
error that returned from the whole function (no ready widget, ZDOTDIR left at
Orca's wrapper dir), and `setopt ksh_arrays` made the 1-based feature
subscript drop whichever feature is listed first.
- The `/etc/zshrc` HISTFILE repair is no longer behind the `history` guard: it
undoes damage Orca's own ZDOTDIR caused, so it must also run for a shell that
re-enters the wrapper after the allowlist was consumed.
- The relay keeps its own wrapping gate. Its .zshenv resolves the user's config
dir from a ZDOTDIR Orca has already overwritten, so wrapping a remote pane
just for `history` cost a relocated-ZDOTDIR user their whole shell config.
- A failed primary spawn no longer leaks the primary shell's launch env
(wrapper ZDOTDIR + feature channel) into an unwrapped fallback pane.
* fix(shell): drop the relay wrapping gate and stop HISTFILE inheriting across Orca instances
The relay-specific gate added last round rested on a false premise:
main's hasOverlayRestoreEnv already included ORCA_REMOTE_CLI_BIN_DIR, and
ssh-pty-spawn-env sets that on every SSH pane whose session has a CLI
bridge — so ordinary remote zsh panes were already wrapped. The gate only
bit where remoteCliBridgeEnv is null (a host too old to report its
platform), where it silently dropped that pane's worktree history. All
three transports now share the features.length rule.
HISTFILE stays exported, so a newly wrapped pane handed the worktree
history path to every child, including a nested Orca whose panes then all
hit injectHistoryEnv's check-before-set and appended into the launching
worktree's file. Same class as the fish_history fix in #15195: recognise a
path Orca minted and drop it before the check, on the desktop, daemon and
relay injection paths and both history-disabled branches.
Also: run the epilogue from the wrapper .zshrc when zsh is in sh/ksh
emulation, since sourcehome() then reads $HOME/.zlogin and the wrapper's
.zlogin never runs; track fallback launch-env keys per attempt rather than
once from the primary; and guard the cross-file epilogue call so a wrapper
dir shared by two builds degrades quietly.
* fix(shell): make every wrapper file self-sufficient and retire the deleted marker vars from tests
- .zprofile/.zshrc/.zlogin each define __orca_resolve_user_config_dir. They
called it on line 2 while only .zshenv defined it, so a wrapper dir written by
two concurrently installed builds printed three "command not found" and
skipped the user's entire zsh config. New live-shell test covers it.
- Retarget every remaining ORCA_SHELL_READY_MARKER/ORCA_SHELL_STARTUP_IDENTITY
reference onto ORCA_SHELL_FEATURES, or delete it where the key is now dead.
- isOrcaMintedHistFile requires a leading '/', so a relative path of the same
shape stays the user's.
- Drop an unused no-control-regex disable, and register the two real-zsh suites
in the dedicated shell-contracts lane.
* fix(shell): stop the zsh wrapper colliding on REPLY and degrade under sh emulation
Widening wrapping from overlay/startup panes to every zsh pane turned three
latent wrapper defects into user-visible ones.
- The config-dir resolver used `REPLY`, zsh's shared scratch global, as its
out-parameter. `typeset -r REPLY` in a user config made the wrapper's first
executable assignment fatal, `typeset -i REPLY` silently resolved every path
to 0; both left HISTFILE inside Orca's wrapper dir. It now writes an
Orca-private `_orca_resolved_config_dir`, declared `typeset -g` so the
rename introduces no `warn_create_global` noise. A new rule test fails on any
generated wrapper file that writes a global outside Orca's namespace.
- The daemon dropped an inherited HISTFILE but never an inherited
ORCA_HISTFILE, which now both wraps a pane the client scoped nothing for and
re-exports another worktree's history path. The relay had the same gap on its
isolation-off and revive paths. Both now mirror the desktop.
- A user .zshenv or .zprofile ending in `emulate sh` makes zsh ignore ZDOTDIR,
so no later wrapper file is read and the epilogue never runs. Nothing can
repair HISTFILE from there, so the wrapper now detects the emulation and
hands the pane back unwrapped instead of leaving history somewhere invisible.
Also `typeset -g __orca_in_command` so the OSC 133 preexec hook prints no
warning under `setopt warn_create_global`.
|
||
|
|
4b2c901b66 |
test(terminal): pin that the CJK block is the preedit overlay, not the cursor (#15242)
* test(terminal): pin that the CJK block is the preedit overlay, not the cursor A report described the cursor sitting on a wide character's first cell and hiding its right half, with cursor style and opacity settings ignored. Neither defect reproduces. Replaying the reporter's own captured byte stream leaves the cursor at column 11, exactly where the application asked, with correct wide and continuation cells. A block cursor also cannot hide half a glyph: it inverts the cell and the syllable renders inside the cursor span. The black block is the IME preedit overlay. macOS 2-set Korean keeps the trailing syllable composing until a terminator, so it sits in an opaque absolutely-positioned box over the grid rather than in the buffer. That box took stock upstream colours, black on white. It explains what no cursor theory can: the block appears at the composing cursor cell, no cursor option reaches it, it is identical with GPU acceleration off since it is a DOM node above both renderers, Latin never triggers it because Latin opens no composition, and Enter clears it because Enter commits the composition. Already fixed by the overlay theming in #15014, which landed a day after the reported release, so the fix ships in the next one. Tests only, no production change. Two pin the negative results so the cursor explanation cannot be re-derived, and one pins the actual mechanism at end of row, beside the existing mid-line arm. Separately confirmed and not fixed here: the WebGL renderer drops the cursor colour's alpha, so terminal cursor opacity genuinely does nothing for a block cursor, which is the default style on the default renderer. That is in the webgl addon rather than in xterm or in our code. Refs #12729 * test(terminal): make the cursor precedence assertion real and measure the overlay Review of the first pass found one assertion that could not fail. It set options.cursorStyle and then read decPrivateModes.cursorStyle, which are separate fields with separate storage, so it pinned that writing one does not clobber the other. Deleting the precedence expression from both renderers left it green. It now asserts the rendered cursor class: the option style renders, a DECSCUSR overrides it, and the reset hands control back. That fails if the precedence is removed. The overlay's rendered width is the one measurement in the report that argues against our explanation, and no test here could reach it, because the unit environment performs no layout. Adds an end-of-row browser arm beside the existing mid-line one, asserting a single composing Hangul syllable spans about two cells. That settles whether the block the reporter measured at one cell can be this overlay. Also scopes two DOM queries to the test container rather than the document, and attaches the render listener before writing so a missed render fails instead of hanging to timeout. Records in the file header what it does not establish: composing the opacity into the theme is not the same as it reaching the screen, since the webgl renderer drops the cursor colour's alpha for a block cursor. Refs #12729 |
||
|
|
442b46f020 |
ci(e2e): install a CJK font on the e2e runners (#15259)
The runners have no font covering Hangul, Han or Kana, so any spec that asserts how CJK text renders is measuring tofu rather than the glyph. That is not hypothetical. An end-of-row preedit spec added for #12729 measured the composition overlay at 1.02 cells against an 8.43px grid and failed its "wider than one cell" assertion. The overlay is shrink-to-fit with no width of its own, so it tracks the glyph's advance rather than the two cells the grid reserves for a wide character. With no Korean font that advance is one cell, and the assertion cannot distinguish a real result from a missing font - which is exactly the question that spec exists to answer. fonts-noto-cjk covers all three scripts and is added to both jobs that execute specs, the sharded suite and the changed-spec job, plus the ssh docker lane so the three stay consistent. This does not make any spec pass on its own. It makes the CJK ones mean something. |
||
|
|
7a695c70f1 |
test(e2e): harden triaged CI failures (#14656)
* test(e2e): harden triaged failures
* test(e2e): ship relay bundle to reusable shards
* test(e2e): tolerate expected IPC closures in daemon shutdown
A normal client exit can close the IPC channel before the finish ack
lands. Distinguish this from real failures by checking error codes,
only throwing if forced cleanup occurred or the error is not an IPC
closure.
* rm doc
* test(e2e): return termination status from legacy close handler
- terminateLegacyCloseClient now returns a discriminated union indicating
whether the process had already exited ('already-exited') or termination
was actually attempted ('termination-attempted')
- Allows finishLegacyCloseClient to only set forcedCleanup when termination
was genuinely needed, not when the process exited cleanly on its own
* test(e2e): fix dispatch contract and voice mic locator
Point the release E2E contract at the renamed build step, and assert the
relabeled microphone through the Voice pane combobox even when Radix
leaves the listbox open.
* test(e2e): add contract test for relay artifact dispatch
Validate that the relay artifact built in CI is properly uploaded,
downloaded, and passed via ORCA_RELAY_PATH to E2E test runs.
* Distinguish between terminated and already-exited processes
Detect when processes have already exited instead of always reporting
termination success. Return booleans from cleanup functions to indicate
whether they actually signalled a process, catch tree-capture failures
when the root process exits before recording completes, and use these
signals to return accurate exit status from termination handlers.
* test(e2e): stabilize file creation and voice microphone tests
Use stable locators (aria-autocomplete, named triggers) and add retry
logic to handle file scans and device events that can interfere with
listbox state. Increase timeouts to allow async operations to complete.
* Add retry logic for transient GitHub API errors in PR body updates
GitHub API occasionally returns transient 5xx errors. Retry up to 3 times
with exponential backoff (1s, 2s, 4s) to improve reliability during
temporary service disruptions. Export updatePullRequest and add sleepImpl
parameter for test injection.
* Add tab search result retention during typing
Keep search results on screen while the deferred query catches up with
the live query. Re-validates results against the current input without
dropping rows prematurely, ensuring the user can select from what they see.
* Add proper types to tab search mock
Replace `unknown` with concrete types (`OpenTabSearchResult`,
`OpenTabSearchEntries`, `SearchableWorkspaceTab`) and use type guards
for discriminated unions to improve test type safety.
|
||
|
|
41ab3b825a |
ci: run the IME e2e suite when terminal input code changes (#15239)
PR e2e only runs specs whose own spec file changed, plus two explicit source-to-spec mappings. Neither covers the terminal pane or the xterm patch, so every terminal IME fix in the 1.4.18x window shipped without triggering a single e2e spec - including one whose diagnosis was later refuted by hardware, and one that turned out to fix a different bug than it claimed. The suite it skipped is not thin. It drives real compositions over CDP, reads real pty bytes, asserts real overlay geometry, and exercises the macOS-only input path on Linux runners through a user-agent policy override. It simply was not pointed at the code it covers. The SSH block immediately above records the same lesson from the same cause: "the Docker-SSH specs only ever ran when someone edited a spec, so four pane-restore regressions shipped from SSH source edits that touched no test." This applies it to terminal input. config/patches is included because the terminal's composition and key handling live in the xterm patch, so a change there is exactly the kind this suite exists to catch. Verified by replaying the filter against the merges that skipped it: 15218, 15223 and 15198 each now select six specs. |
||
|
|
49752477a6 |
build(xterm): restore the patch regeneration harness and gate it in CI (#15223)
* build(xterm): restore the patch regeneration harness and gate it in CI docs/reference/ime-architecture.md says "Never hand-edit the bundles in the patch" and links to docs/reference/xterm-patch-regeneration.md. That doc does not exist, and neither does the harness it describes. Both landed in |
||
|
|
e2b567363b | ci: stop refreshing every apt repo three times to install fish (#15217) | ||
|
|
13b10e0b54 |
ci: cut PR wall clock by caching what CI recomputes every run (#15211)
None of these change what CI checks — they remove work the runners repeated on every PR. - install-node-dependencies installed with --no-frozen-lockfile, so every job re-resolved the graph against the registry to recompute what the lockfile already pins. Measured at ~62 MB of packument metadata per job; the pnpm store cache does not cover the metadata cache, so this was paid ~39 times per run. The `git diff` guard that made the re-resolution redundant stays. - --ignore-scripts leaves node-pty with no build/Release, so ensure-native-runtime node-gyp-compiled it in every job asking for a runtime. Cache the build under an ABI-bound key (runtime, resolved Node version, node-pty patch) with no restore-keys, since a partial match is exactly the mismatched build that would be recompiled anyway. - The four fetch-depth: 0 checkouts pulled full history including every historical blob (blobs are ~89% of this repo's pack). They only need the commit graph for a merge-base diff, so fetch them blobless. Measured 30-43s each today versus 8s for the shallow checkouts. One of them, e2e-paths, gates the entire E2E chain. - E2E jobs ordered setup-node before pnpm, which meant setup-node could not find the store and no E2E job cached dependencies at all. Reorder and cache; this sits on the critical path in both the build job and each shard. - git_compatibility rebuilt Git 2.25.5 from a pinned tarball on every PR. Cache the build; the sha256 assertion still guards the miss path. - typecheck ran three independent tsc passes back to back and discarded the .tsbuildinfo each project already emits. Run them concurrently and cache the incremental state. - package (windows) built the electron-vite targets serially via build:release. Use a :parallel variant that overlaps them, matching what the Linux package job already packages and smoke-tests from. Contract tests cover each new cache's ordering and key so none of them can silently start serving a stale or ABI-mismatched artifact. |
||
|
|
24e662adc1 |
feat(ssh): verify host keys, and restore panes correctly across a reconnect (#14844)
* docs(ssh): design for real host key verification (STA-4319)
Today's ssh2 verifier records a fingerprint and returns true — every host key is
accepted, with no known_hosts consult and no change detection anywhere in
src/main/ssh/. Scope is per-connection, so exec, SFTP, port forwarding, the
watcher and relay deploy all ride that one unverified handshake, and the
ProxyJump path puts the final hop — the topology most likely to cross untrusted
network — on ssh2 specifically.
Decisions worth calling out:
- Read the user's known_hosts as a trust source but NEVER write to it. That file
is shared with every other SSH tool on the machine; appending means line
endings, permissions, concurrent writers and a corruption blast radius well
beyond us. Accepted keys go to our own per-target store. Reading theirs is also
the entire migration story: most developers already have their hosts there.
- Mismatch is scoped to the SAME key type. A host with only an RSA entry that
presents ed25519 is unknown, not changed. ssh2 negotiates ed25519 first, so
without this we would fire a change-of-key alarm at nearly every existing user
on their first upgraded connect — training them to dismiss the one warning that
is supposed to mean something. Flagged in review as the decision I am least
sure of; a downgrade-vector argument against it is being tested.
- Changed key hard-fails with no override button; recovery is a separate explicit
action, offered only when OUR store is what disagreed, because forgetting our
record cannot unblock a known_hosts conflict.
- Background reconnects deny rather than prompt. A dialog the user cannot place
in context only teaches click-through.
Two traps are documented because either would make the fix silently do nothing:
an async verifier returns a Promise, which ssh2 reads as truthy and accepts
immediately; and the existing test mock invokes hostVerifier with one argument
and ignores the return, so it would pass against a verifier that never decides.
Design only — no behaviour change. The doc is added to the tracked-reference
allowlist in .gitignore alongside the other docs/reference entries.
* docs(ssh): revise the host key design after security and migration review
Three things the reviews changed, kept visible rather than quietly edited out.
THREAT MODEL WAS WRONG IN THREE PLACES. Jump hosts are not the worst case — they
are already safe: shouldUseSystemSshTransport branches on exactly the inputs
resolveEffectiveProxy does, and attemptConnect returns after the system probe, so
ProxyJump goes through OpenSSH and is verified. Agent forwarding was overstated
(gated on the user's ForwardAgent). Credential theft was understated: any auth
error counts as agent fallback, so a MITM walks the user to the password AND
private-key passphrase prompts, and cachedPassword replays without prompting. The
relay claim was backwards — the attacker owns their own machine; the real impact
is the return direction, where they become the host our workspace trusts.
TYPE SCOPING IS A DOWNGRADE VECTOR WITHOUT ALGORITHM ORDERING. This was the
decision I flagged as least certain and asked to have argued both ways. OpenSSH
is safe only because order_hostkeyalgs() puts known types first and RFC 4253
gives the client's order priority. ssh2 negotiates ed25519 first regardless, so
an attacker who cannot forge the RSA key on file just presents ed25519 and gets a
friendly first-contact prompt instead of a hard failure. Keep scoping, but set
algorithms.serverHostKey to lead with the types on file — and add a sixth
outcome for 'unknown type, known host', which must never read as first contact.
SHIP THE DEFENCE BEFORE THE DIALOG. Startup restore fires eager connects for all
targets in parallel with a 15s timeout while a prompt would live 120s; ephemeral
VM targets present a new key every launch; paired-web connects run on the host
desktop, so the dialog opens on someone else's screen. Phase 1 is therefore no
modal at all: consult known_hosts and our store, match connects, unknown persists
with accept-new semantics, mismatch and revoked hard-fail. That is the whole MITM
defence with none of the migration risk.
Also folded in, verified live against OpenSSH 10.2p1: the without-port fallback
(bracketed lookup first, then bare, where the second pass can only yield match or
unknown — otherwise a bare line plus a non-default port produces a spurious
prompt); hashed entries hash the candidate form; multiple files union; a
cert-authority line does not match a plain key. IPv6 and bracket parsing moved
INTO scope — that is a parser requirement, not a scope call, and getting it wrong
produces the prompt-training harm the design exists to avoid.
* feat(ssh): parse and match OpenSSH known_hosts
The matcher half of STA-4319. No behaviour change yet — nothing calls this.
Hand-rolled because no maintained JS implementation exists, and written against
behaviour observed from OpenSSH 10.2p1 rather than inferred from the man page.
Three of those behaviours a reasonable reading gets wrong:
- A non-default port is TWO ordered lookups, not one candidate set: '[host]:port'
first, then bare host ('checking without port identifier' in ssh -v). The
fallback pass can only yield match or unknown — OpenSSH downgrades a wrong key
there rather than reporting a change. Collapse them and anyone holding a bare
line who connects off-port gets a spurious first-contact result; treat the
fallback as authoritative and they get a false change-of-key alarm.
- Revocation resolves in its own pass so the verdict cannot depend on line order.
Verified both orderings.
- A cert-authority line never matches a plain host key; it only validates
certificates. A normal line alongside it still decides.
Mismatch is scoped to the same key type, and a host known by a DIFFERENT type
returns unknown-type-known-host rather than plain unknown — an attacker who
cannot forge the key on file must not get a friendly first-contact result by
presenting another type. That outcome is only half the defence; the other half
(leading serverHostKey with known types) lands with the wiring.
47 tests from vectors executed against real sshd, including ssh-keygen -H hashed
entries. Each of six mutations reddens it: collapsing the passes, letting the
fallback report mismatch, dropping type scoping, resolving revocation in line
order, honouring an unrecognised marker, and skipping the blob/type agreement
check.
* feat(ssh): decide what to do with a presented host key
The policy half of STA-4319, kept separate from the ssh2 wiring so it is testable
without a handshake and injected rather than importing its sources, so a test
states its own trust state instead of writing files.
Phase 1 ships no dialog — a test asserts the decision is never 'prompt'. Startup
restore opens every previously-active target at once, ephemeral VM targets would
ask every launch, and paired-web connects run on the host desktop where the
dialog would appear on someone else's screen.
Ordering that matters: revocation outranks everything including
StrictHostKeyChecking=no, because a revoked key is a statement that this key is
known-bad rather than merely unrecognised. known_hosts is named before our own
store on a change, because its remedy (ssh-keygen -R) is the one that also
unblocks ssh and git — pointing at a remedy that cannot work is worse than none.
Two carve-outs with reasons: an ephemeral runtime target accepts WITHOUT
recording, since a fresh VM presents a new key every launch and a stored record
would accumulate per launch and eventually read as a spurious change; and when
ssh -G ran on the HOME-divergent path that suppresses /etc/ssh/ssh_config, an
unknown host is denied, because a site-wide policy may forbid it and being laxer
than ssh is the one outcome that is never acceptable.
Rejection text deliberately avoids 'authentication failed' and 'permission
denied': the reconnect ladder classifies on those substrings, so a denial phrased
that way is retried forever against a decision that will never change. Pinned by
a test.
* feat(ssh): build the host key verifier and the algorithm order that makes it safe
Still not wired into the handshake — that lands next. This is the piece that
turns a decision into an ssh2 callback, plus the half of the design that is easy
to forget because it lives in a different config field.
The verifier MUST be a plain function returning undefined. ssh2 does
'const ret = verifier(key, verify); if (ret !== undefined) verify(ret)', so an
async function returns a Promise — neither undefined nor falsy — and ssh2 accepts
the key immediately while ignoring whatever the callback later decides. Making
this async would silently restore exactly the accept-everything behaviour the
module exists to remove, so a test asserts the return value is undefined.
orderServerHostKeyAlgorithms is what makes type-scoped matching safe rather than
a downgrade. RFC 4253 gives the client's algorithm order priority, so leading
with the types we already hold for a host denies a server the choice of
presenting some other type to convert a hard failure into first contact. Without
it, an attacker who cannot forge the key on file just offers a different
algorithm. Revoked entries never contribute to that order.
Also fails closed on two paths that would otherwise hang or over-trust: a key
whose own length-prefixed header cannot be read is refused rather than reasoned
about, and a throw from any dependency denies, because ssh2 may not catch an
exception raised inside the verifier and the handshake would hang instead of
failing.
18 tests. Includes the two negative cases that matter — first-contact keys are
recorded, but keys we already know, rejected keys, ephemeral runtime targets and
a lax StrictHostKeyChecking are not.
* fix(ssh): promote every RSA signature algorithm for a known ssh-rsa key
A known_hosts entry names the KEY type, which is not the negotiated ALGORITHM
name. One ssh-rsa key is offered as rsa-sha2-512, rsa-sha2-256 or ssh-rsa
depending on the signature algorithm, so matching the literal name only would
leave a host we know by RSA ordered behind ed25519 — precisely the ordering this
function exists to prevent, and precisely the population (RSA-era known_hosts
entries) it was written for.
Verified from ssh2's own negotiation while wiring this: kex.js iterates the
CLIENT list and takes the first entry the server also offers, so client order
does decide, as RFC 4253 says. ssh2's default order leads with ed25519 and places
the RSA algorithms fifth through seventh.
* fix(ssh): verify host keys instead of accepting every one (STA-4319)
The actual fix. ssh-connection's verifier recorded a fingerprint and returned
true, so every ssh2 connection accepted every host key — no known_hosts consult,
no change detection. It now consults the user's known_hosts plus our own store
and refuses a changed, revoked or unverifiable key.
Phase 1 by design: no dialog. Unknown hosts are accepted and recorded
(accept-new semantics), because startup restore opens every previously-active
target at once, ephemeral VM targets present a new key each launch, and
paired-web connects run on the host desktop where a prompt would appear on
someone else's screen. The MITM defence lands now; the prompt is Phase 2.
Also sets algorithms.serverHostKey to lead with the types already known for the
host. Without it the type-scoped matching is a downgrade — an attacker who cannot
forge the key on file just presents another type and turns a hard failure into
first contact. Verified from ssh2's kex.js that the client list decides.
Denial replaces ssh2's generic handshake error with the specific reason, because
the reconnect ladder cannot distinguish a generic failure from a transient fault
and would retry forever against a decision that will never change.
An unreadable trust store degrades to known_hosts only rather than failing the
connect: a changed key is still refused, and a host trusted only by us falls back
to first contact and is re-recorded, reaching the same decision.
The ssh2 mock now uses the callback form and aborts the handshake on denial. As
written it called hostVerifier(key) with one argument and ignored the result, so
it would have passed against a verifier that never decides — flagged in the
design as a mock that had to change, not a test to quietly rewrite. Two new tests
pin the wiring rather than the module: an unidentifiable blob is refused, and a
well-formed key is accepted.
Note for review: commit
|
||
|
|
7b4e10b104 |
fix(mobile-ios): pin fastlane and gate the Fastfile in CI (#15092)
* fix(mobile-ios): pin fastlane and gate the Fastfile in CI The ios-distribute job failed on every run from 2026-08-10 to 2026-08-13 because distribute_testflight passed distribute_only without app_platform, so pilot fell through to an interactive platform prompt on ubuntu. No CI check loads the Fastfile, so external testers got nothing for six days. - Pin fastlane 2.238.0 and commit mobile/Gemfile.lock so ios-build (macos) and ios-distribute (ubuntu) cannot resolve different versions ~25 minutes apart. Fixes the Gemfile comment's dead mobile-build.yml reference. - Add a Fastfile smoke check (bundle exec fastlane lanes) plus a static contract test for the TestFlight lane arguments to Mobile Checks. - Set reject_build_waiting_for_review so a superseded same-train build in beta review stops blocking the submission. * fix(mobile-ios): install the pinned Gemfile.lock in frozen mode Without frozen, a lockfile that drifts from the Gemfile is silently re-resolved per job, which is the version split the pin exists to prevent. * test(mobile-ios): anchor the TestFlight argument contract against an empty selection * chore(mobile-ios): canonicalize the lockfile platforms Bundler's own normalization drops arm64-darwin-25 as redundant with the versionless arm64-darwin, and the ubuntu runners resolve x86_64-linux-gnu. |
||
|
|
3bb87ff93b |
reland(shell): one portable Unix startup dialect, with both revert causes fixed (#15018)
* reland: portable startup-shell dialect, with the two revert causes fixed Relands #14863 (reverted by #14975) with fixes for both regressions the revert cited. 1. History GC deleted folder-workspace shell history. The live set was built from `getAllWorktreeMeta()` alone, but a folder workspace's PTY carries `folder:<id>` as its worktree id, so every live folder workspace looked orphaned. `getKnownWorktreeIdsForHistoryGc` now unions in `getFolderWorkspaces()`. Both consumers — the history-directory prune and the fish-history sweep — read that one set, so the fix covers bash, zsh and fish history alike. The directory prune had this gap since #1524; #14863 only widened its blast radius to fish files. 2. A copied Codex resume command aborted under `set -u`. Its leading clear statement has to test `$fish_pid`, and that unbound expansion takes the whole line — including the agent launch — down with it. Copied text runs in a shell Orca never spawned, so nothing can seed that variable first. The removal now rides on the agent itself as `env -u`, which needs no shell syntax and no expansion. Verified byte-identical under `set -u` in sh, bash, zsh, dash, ksh and fish. `env` cannot run the `cd` builtin, and a child `cd` would not move the agent, so the prefix is placed on the agent rather than on the whole `cd … && agent` chain. cmd and PowerShell have no nounset hazard and keep their clear ahead of the `cd`, which preserves `cd … && agent` — a failed `cd` still cannot launch the agent in the wrong directory. * fix(history-gc): stop three more paths from deleting live shell history Found by adversarial review of the reland. All three are the same class as the bug that caused the revert: a live set that is missing a category of real workspace, so the GC reads it as orphaned. 1. Profiles. The history root is `userData/terminal-history`, which has no profile segment, but the Store the GC consults is per-profile. So after a profile switch the live set condemned every other profile's history — and fish history, which lands in the user's own fish data dir, is shared by every profile on the machine. The live set now unions in the inactive profiles' worktrees and folder workspaces, read from their data files. A profile whose ids cannot be read reports the empty set rather than one that condemns real history. 2. No empty-set guard on the tree scan. `sweepOrphanedFishHistoryFiles` refuses an empty live set because it cannot be told apart from a store that failed to hydrate; the directory scan, which deletes more, had no such guard. A store that fell back to default state would have taken every worktree's bash and zsh history with it, across all roots including WSL. Four existing tests passed `new Set()` and relied on "empty means everything is orphaned" — exactly the behavior being removed — so they now pass a real live set. 3. Relay fish history. The relay isolates its history tree under its own root but wrote fish history into the shared fish data dir under the desktop naming, keyed by the CLIENT's worktree ids. On a machine running both Orca and a relay host, the desktop sweep deleted remote sessions' history once it went stale. Relay files are now `orca_relay_<hash>`, which the sweep's pattern deliberately does not match; the relay still deletes them by exact name when the worktree goes away. * fix(resume): enforce the env-removal invariants instead of documenting them Both found by adversarial review; both were unreachable from today's callers and silent if reached, which is exactly how they would survive to a caller that does reach them. - A pinned CODEX_HOME and the removal named the same variable, and `env -u` strips what the assignment just set — so the agent would have resumed against the real home and not found the session. The removal list now excludes any name the prefix pins, keeping the assignment authoritative as the old `clear…; CODEX_HOME=x agent` ordering did. Same fix in the git-bash twin. The PowerShell branch already clears before it assigns, so it was never affected. - Placement was keyed on the platform while the grammar it selects is keyed on the shell, so `platform: 'linux'` with `shell: 'powershell'` emitted POSIX `env -u` into a PowerShell line. PowerShell now routes to the PowerShell builder whatever the host, and the POSIX/cmd split below asks the shell rather than the platform. |
||
|
|
08bf209e40 |
fix(ci): run PR LoC scripts from the default branch, not PR head (#15016)
The PR test LoC job fetched .github/scripts/pr-test-loc-*.mjs from pull/<n>/head and ran them with node while holding a GITHUB_TOKEN scoped pull-requests: write, so PR-authored code executed under a write token. Pin the fetch to the repository default branch. base.sha is not enough: for stacked PRs it is an unreviewed feature-branch commit any collaborator can push to, while main is gated by branch protection. Also pass event data via env instead of shell interpolation, and add set -euo pipefail so a failed download cannot leave a truncated script. |
||
|
|
f070033156 |
Revert "refactor(shell): one portable Unix startup dialect instead of shell d…" (#14975)
This reverts commit
|
||
|
|
b6ea3f17a9 |
refactor(shell): one portable Unix startup dialect instead of shell detection (#14863)
Orca had to guess which shell would parse a queued command line, then emit syntax for it. Guessing is unreliable for a remote or WSL host, and every dialect-dependent function is a place to get it wrong. Replace the guess. Everything emitted for a Unix shell is now built to be correct in sh, bash, zsh, dash, ksh and fish alike, so no detection is needed: - quoteStartupArg emits backslashes as "\\" and apostrophes as "'" between single-quoted runs. Both families read that identically, unlike the sh '\'' idiom, which fish silently halves and which makes a trailing backslash a hard syntax error. - clearEnvCommand emits a self-contained fish/sh branch. It deliberately does NOT call a helper defined by Orca's shell wrappers: Orca wraps only zsh, bash and fish, so an `sh`/`dash`/`ksh` login shell launches unwrapped — and the same text is copied to the clipboard and pasted into shells Orca never spawned. In both, a helper would be `command not found`, which is the exact failure this exists to avoid. Two guarded statements rather than `A && B || C`, because fish's `set -e` returns non-zero for an already-unset variable and would fall through to the sh branch; a trailing `true` pins the status, since this is the last statement of a launch line and the prompt renders it. - One tokenizer for Unix. The input is a settings string the shell never parses, so parsing it per-shell only made the same setting mean different things in different workspaces. AgentStartupShell loses its 'fish' and 'unix' members, and the three login-shell resolvers, the fish tokenizer and the agentEnv.SHELL probe go with them. Per-worktree shell history now actually works: - zsh on macOS was a no-op. /etc/zshrc assigns HISTFILE unconditionally before any wrapper Orca controls, so the injected value was already gone — and with ZDOTDIR still pointing at Orca's wrapper dir, history landed inside it. The intended path rides ORCA_HISTFILE and is restored after user config. Fixes #11044. - fish keeps history in its own data dir keyed by session name, since it ignores HISTFILE and has no custom-directory knob. Files are deleted rather than truncated, a symlinked ~/.local/share no longer disables cleanup, and a GC sweep reclaims orphans whose meta.json is gone. The sweep refuses an empty live-worktree set (indistinguishable from a store that failed to hydrate) and skips files younger than GC_MIN_AGE_MS, mirroring the tree GC's guard against the live-set snapshot race. Verified against real shells rather than asserted as strings: startup-shell-portability.live-shell.test.ts runs 194 assertions across sh/bash/zsh/dash/ksh/fish, and zsh-scoped-histfile.live-shell.test.ts drives a real login zsh through /etc/zshrc. Both are vacuity-checked. The same quoting corpus was replayed byte-exact on Linux, where /bin/sh is dash. |
||
|
|
fa9b20cb41 | feat(skills): reland private bundle sharing safely (#14934) | ||
|
|
763b1febeb |
Revert "feat(skills): add private bundle sharing (#14401)" (#14913)
This reverts commit
|
||
|
|
757fae28d7 |
feat(skills): add private bundle sharing (#14401)
Co-authored-by: E2E Test <e2e@test.local> |
||
|
|
1b6d2403cb |
ci: run full e2e against the daily cut commit (#14870)
After a live daily publish, dispatch e2e.yml at the cut SHA. Detached on purpose so a red suite cannot fail or delay the signed daily. |
||
|
|
5c56bfb28b |
ci: run the daily macOS build 4 hours later (#14869)
The 14:15 UTC cut is too early (6:15am PST / 7:15am PDT). Move it to 18:15 UTC so dailies land late morning Pacific instead. |
||
|
|
393c8764e0 | ci: post test vs non-test LoC on pull requests (#14738) | ||
|
|
9367169888 |
refactor(tests): split every oversized test file off the max-lines suppression list (#14728)
* refactor(tests): split oversized test files off the max-lines suppression list Every `*.test.ts`/`*.spec.ts` that carried an `eslint/oxlint-disable max-lines` directive is now split into focused, behavior-scoped suites that fit the 800-line test budget, with shared setup extracted into co-located `*-test-harness.ts` / `*-test-fixtures.ts` modules (300-line budget). 83 files became ~930; the largest output is 797 effective lines. `orca-runtime.test.ts` is intentionally untouched. Test bodies were moved by scripted line-range slicing rather than retyped, so assertions are byte-identical. The only permitted body edits were mechanical rebinding where a shared value moved into a harness (e.g. `tmpHome` -> `homes.tmpHome`). Registries that enumerate test files were updated in lockstep: - config/max-lines-baseline.txt: pruned 341 -> 258 entries (all 83 removed). - config/reliability-gates.jsonc: 33 gates repointed at the split files, with assertionRefs split per file where a gate's coverage now spans several. - .github/workflows/pr.yml: the real-zsh lane now lists the 4 split files that actually exercise zsh, so they keep running in the dedicated shell lane. Also renamed agent-hooks `server-test-fixtures.ts` to `server.test-fixtures.ts` so the global-fetch call-site audit keeps skipping it, and added `.js` extensions to the CLI suites' dynamic harness imports (node16 resolution) to unbreak `build:cli`. Verification: full suite 52,449 passing vs 52,448 at baseline with zero assertions lost; `pnpm lint`, `pnpm typecheck`, and `pnpm build:cli` all exit 0; the terminal-pane e2e spec runs 31/31 headless. * refactor(tests): split hook-idle arbitration suite that oxfmt pushed over budget The pre-commit oxfmt pass reflowed pty-connection-hook-idle-arbitration.test.ts to 811 effective lines, 11 over the test budget. Split the hook-completion side effect and replacement-agent veto cases into their own suite; both files now sit well under the cap and the 15 tests are unchanged. * test: port upstream test changes into the split files after rebase Rebasing onto main surfaced 27 tests that main had added to files this branch deleted, plus edits to tests that had already moved. Taking the deletion side of those modify/delete conflicts would have dropped that coverage silently, so each upstream change is ported into the split file that now owns the behavior — for example main's six orchestration mailbox tests land across orchestration-runs, -send, and -check. Also repoints `orchestration.notification-mailbox-consistency`, a gate main added after this branch's gate remap, at those same three split files, and re-prunes the max-lines baseline against main's (257 entries). Verified: all 27 upstream test titles present; full suite 52,761 passing with the only diff vs baseline being 12 tests main itself removed and 3 that moved from skipped to passing; lint and typecheck exit 0. * fix(test): flush pending continuations before tearing down terminal test globals CI shard 5/16 failed on both Node 24 and 26 with `ReferenceError: window is not defined` from pty-connection.ts, surfacing through pty-connection-daemon-snapshot-replay.test.ts. The reattach/settle chains `await` a real promise and then touch `window.api`. Under fake timers those continuations cannot run, so they only become schedulable once restoreTerminalTestGlobals() switches back to real timers — which previously happened immediately before `delete globalThis.window`, so a late continuation threw and failed the whole file. Flush async ticks in that window instead. This is latent in the source rather than new: the pre-split 25k-line file kept running other tests after these, which gave the chains time to settle before teardown. Splitting the file moved teardown directly behind them. * fix(test): keep an inert window after terminal test teardown instead of deleting it The async-tick flush was not enough: the reattach/settle chain can resolve after teardown regardless of how long we drain, so CI shard 5/16 still failed with `ReferenceError: window is not defined` from pty-connection.ts. A real renderer never loses `window`, so deleting it was the artificial part. Swap in an inert proxy whose properties resolve to callables and whose calls resolve to undefined, making a late `window.api.pty.*` call a harmless no-op. The next test replaces it wholesale via installTerminalTestGlobals(), and no test asserts that `window` is absent. |
||
|
|
8b22f044f5 |
fix(vm): preserve runtime sidecar rollback compatibility (#14444)
* test(vm): reproduce runtime store rollback poisoning * fix(vm): keep runtime sidecar rollback-readable * fix(vm): harden rollback-compatible runtime persistence * fix(vm): publish rollback lifecycle authority first * test(vm): harden rollback compatibility coverage |
||
|
|
cbed44410a |
STA-4276: preflight Codex in Command Prompt and Git Bash (#14441)
* fix(terminal): preflight Codex in Windows cmd and Git Bash * test(terminal): run Windows preflight through ConPTY * test(terminal): isolate cmd harness exit status * test(terminal): allow slow Git Bash ConPTY startup |
||
|
|
cd6114ab7e | fix(browser): acknowledge paired tab before navigation (#14402) | ||
|
|
94df72d9eb |
ci(windows): cover the worktree admin fingerprint on the Windows runner (#14378)
The fingerprint gate added in #14207 reads Git's administrative layout directly -- `.git` as a file or directory, `commondir`, and per-worktree `HEAD`, `gitdir`, and `locked` -- instead of shelling out to `git worktree list`. That makes it depend on Windows path resolution, CRLF inside those files, and whether `worktree move`/`lock` and deleting a live checkout behave as they do on POSIX. PR CI runs the vitest suite on ubuntu-latest only, so none of that was exercised. Both suites were verified by hand on a real Windows host (Git 2.55.0.windows.3, Node 24.18.0) and pass 25/25, but nothing kept them passing. Add them to the existing curated `Test Windows-specific boundaries` step rather than standing up a new job: the `package (windows)` job already checks out and installs dependencies, so this costs only the tests themselves. |
||
|
|
77b37d85e2 |
feat(vm): create workspaces from provisioned SSH roots (#14359)
* feat(vm): use recipe-provisioned SSH roots * fix(vm): preserve ordinary create failure timing * test(vm): prepare provisioned root SSH fixture * ci(vm): enable SSH setup for provisioned root E2E |
||
|
|
ff8dda81e8 |
fix(serve): exit cleanly after headless Linux signals (#14334)
* fix(serve): keep owned Xvfb alive through Electron teardown * test(serve): gate packaged signal shutdown * test: harden headless shutdown lifecycle gate * fix(serve): isolate Xvfb from foreground signals * docs(serve): preserve Xvfb during systemd stop * test(serve): pin shutdown policy to owned Xvfb unit * test(serve): harden shutdown gate portability * test(serve): bound systemd unit parsing |
||
|
|
2f41c286e2 |
fix(docs): replace stale preload typecheck reference (#14298)
* docs: fix stale preload typecheck reference Signed-off-by: HoonDongKang <d159123@naver.com> * docs: keep .d.ts guidance canonical --------- Signed-off-by: HoonDongKang <d159123@naver.com> Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> |
||
|
|
70abf5cacc |
test: add golden e2e tests for agent TUI launch and shell recovery (#14258)
* test: add golden e2e tests for agent TUI launch and shell recovery
Add test fixtures and E2E tests to verify agent TUI functionality:
- Stub agent implementation supports cross-platform execution (Unix/Windows)
- Test verifies multiline composer with Shift+Enter support in agent TUI
- Test verifies clean shell resumes after agent exit without state leakage
* test: add golden e2e tests for agent TUI launch and shell recovery
Add agent TUI launch and shell-recovery tests to the golden (release-blocking)
E2E suite, covering agent initialization and shell availability after agent
exit. Improve escape sequence handling in the stub agent to prevent stray key
reports from contaminating test output. Add terminal input readiness checks to
ensure commands execute reliably before verification.
* test: coerce golden stub stdin chunks for type-aware lint
Node types the stdin data event as string | Buffer even after
setEncoding('utf8'), so restrict-plus-operands failed CI.
* test: fix golden stub agent Windows batch files and add Ctrl+C support
- Store batch files with CRLF to avoid Windows 512-byte parser boundary bug
- Handle Ctrl+C (0x03) in raw mode as alternative to Ctrl+D (0x04)
- Update release notes documenting golden test skip behavior on older tags
* Remove Windows batch file gitattributes workaround
The -text whitespace=cr-at-eol rule preventing CRLF conversion for
.cmd files is no longer needed. Allow batch files to use normalized
line endings.
|
||
|
|
b8d6b21dfa |
test(e2e): add golden E2E tests for workspace session management (#14304)
* test(e2e): add golden E2E tests for workspace session management - Restore exact file and terminal state after quit/relaunch - Verify terminal file link activation and external edit detection - Test worktree creation and switching with isolated terminals - Isolate test repo paths between concurrent CI runs with UUIDs * Add platform-aware marker echo command utility - Create splitMarkerEchoCommand() to generate shell commands that safely echo test markers across Windows and Unix platforms - Split markers into prefix/suffix fragments so output assertions prove execution, not just shell echo-back - Consolidate SORTABLE_TAB export and improve tab bar locator logic - Refactor terminal link helpers to extract client point calculation |
||
|
|
e84fb46eaf |
test: add golden E2E tests for source control workflows (#14260)
* test: add golden E2E tests for source control workflows - Tests core source control interactions: file edit/save, commit staging, and diff viewing - Integrated into CI/CD pipelines for Linux, macOS, and Windows - Includes helper utilities for test setup and worktree management * test(e2e): verify golden commit author and fix test flakiness - Configure git author name/email at worktree level during setup - Verify commits are made with correct author details in assertions - Add explicit timeouts to file visibility waits and git status polling - Fix test ordering to seed edits after source control is open - Simplify git status refresh logic to rely on automatic updates * Add rollback to createGoldenWorktree on setup failure Cleanup callbacks only register after setup succeeds. When a config command fails, the half-built worktree and branch leak into later test runs, causing flakiness. Now we roll back immediately and re-throw the setup error. * test(e2e): match explorer rows after the git status badge appears The golden file-save spec used an exact /^README.md$/ filter. After save, the explorer row text becomes "README.md M", so reopen clicked nothing. * test: strengthen golden worktree setup verification - Track working directory in git call inspection to verify correct execution context - Verify user.name/email config applies to worktree-specific settings, not repo - Add exhaustive setup call sequence assertions to catch setup/rollback leaks |
||
|
|
e7b85266f5 |
Add golden E2E tests for fresh terminal and shell commands (#14302)
* test(e2e): add golden tests for fresh terminal and shell commands Adds regression tests for terminal initialization in fresh profiles and shell command execution to the golden test suite, integrated across Linux, macOS, and Windows CI. * test(e2e): bracket shell command output between markers The echoed command can wrap or be clipped by the buffer tail. Bracket output between begin and end markers to reliably identify real output, and strip ANSI escape sequences that interfere with parsing. |
||
|
|
dd63d35d09 |
fix(ci): skip missing golden scripts on older release tags (#14267)
Cut Release is dispatched from main but checks out the tagged tree. Cherry-pick tags such as v1.4.182-rc.1 do not define test:e2e:windows-fresh-startup-golden, so the Windows golden job failed with ERR_PNPM_NO_SCRIPT. Run tag-optional goldens with --if-present. |
||
|
|
7c93aed6dc |
Fix fsync of read-only files on POSIX (#14235)
* fix(files): fsync read-only files on POSIX * test(e2e): add golden E2E tests for POSIX profile index fsync Validates that profile index files are properly persisted on POSIX systems, including with restrictive umask settings. These are release-blocking golden tests for Linux and macOS. * test(terminal): wait for fish child ownership before stdin write Fish 4.8 withdraws DECSET 2031 before spawning the child, so the shell-contracts harness could send hello into an intermediate prompt and hang waiting for CHILD-READ. Wait for the child's CHILD-READY marker and answer split DA1/CPR/OSC queries across chunk boundaries. * test(e2e): verify profile index persists to disk with restrictive umask Strengthen the POSIX fsync test to verify the rebuilt index is actually written to disk and has correct permissions under a restrictive umask, not just cached in memory. |
||
|
|
9a10561258 | fix(terminal): retain SSH startup delivery through reconnect (#14161) | ||
|
|
de729d6067 | Fix paired web creation failure handling (STA-4024, STA-4025, STA-4063) (#14100) | ||
|
|
b115f8d256 |
Add Windows golden E2E test for fresh-startup regression
Windows terminal rendering golden is flaky on CI runners. Re-enable Windows in the golden E2E gate with a scoped test for the fresh-profile startup regression from #14130. Terminal rendering continues on Linux and macOS; Windows runs fresh-startup only. |
||
|
|
e8044b1b30 |
fix(windows): restore fresh-profile startup after durable fsync (#14173)
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> Co-authored-by: DHTheOne <238933622+DHTheOne@users.noreply.github.com> Co-authored-by: Anton Tupitsyn <70199858+PLUTONYY@users.noreply.github.com> Co-authored-by: 7loop <48346764+7loop@users.noreply.github.com> |
||
|
|
90b8554fc9 | fix(terminal): recover readiness after startup exec (#14027) | ||
|
|
5ea7df1a5b |
fix(terminal): make DECSET 2031 subscriptions silent (#13904)
fish arms `CSI ?2031h` before painting each prompt and withdraws it when it hands the tty to a child — a ~1ms window. Orca answered that subscribe with `CSI ?997;Nn` across a 1-3ms renderer hop, so the reply landed after the withdrawal and was read as stdin by the next child, corrupting `brew`/`npx` `[y/N]` prompts. The reply is not stale by Orca's own view when written (measured staleReplies: 0), so no suppress-the-stale-reply scheme can close this — the information needed to suppress does not exist yet. Nothing asked for the reply either. The Contour spec says a terminal "should only send out the DSR when the palette has been updated"; Ghostty (Termio.zig:729 — force=true reachable only from the ?996n DSR), iTerm2 (VT100Terminal.m:995 — flag only) and xterm.js (InputHandler.ts:2035 — flag only) all emit nothing on the DECSET. So stop entering the race: record the subscription, answer nothing. Of 17 real programs measured under a pty, only fish, tmux, claude and opencode subscribe; none block on a reply, and answering produces one redundant palette re-query and zero rendering difference. tmux is the only one that sends `?996n`, which Orca still answers. - Subscribes are record-only at all four emitters (live scan, hidden-gate fact, parked byte watcher, parked responder — the last is deleted, it only replied). - `?996n` answers, the subscription registry, and the theme-flip push are unchanged. `paneLastThemeMode` is still seeded at subscribe so the next appearance re-apply is not read as a flip. - Replay grammar carries `?2031l` alongside `?2031h`, so a late-attaching remote client no longer registers a subscription the TUI already retired. Also closes fish-integration gaps found alongside: `unset` (which fish lacks) becomes `set -e` on paths parsed by the client's login shell, `config.fish` is parsed for agent-home detection, and bracketed-paste startup delivery is made consistent across local/daemon/relay. Regression test drives real fish 4.7.1 under node-pty and asserts on what the child process reads; it fails against pre-fix code with the exact payload from the issue. CI installs fish 4 and fails loudly rather than skipping. Closes #9993 Co-authored-by: Orca <help@stably.ai> |
||
|
|
d6e1d84235 |
fix(wsl): forward native CLI arguments losslessly (#12582)
* fix(cli): preserve WSL --deps quotes and parse task ids strictly PowerShell 5.1 native splat was stripping ASCII double quotes on the WSL bridge, so non-empty JSON --deps arrays failed while [] still worked. Pre-escape quotes before launching orca.exe, and recover quote-stripped task-id arrays while rejecting non-task-id and malformed CSV input (#12188). * fix(orchestration): narrow WSL deps recovery * fix(wsl): forward native CLI arguments losslessly * ci(windows): exercise WSL PowerShell argv boundary --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
4c69e45552 |
Strengthen plain-node-entry-guard with entry name validation (#12761)
* Strengthen plain-node-entry-guard with entry name validation - Add buildStart hook to validate guarded entry names exist in rollup inputs, preventing stale names from silently stopping guards - Extend electron require detection to subpaths (electron/main, etc) - Improve smoke test signal/exit handling and use constants - Add comprehensive tests for entry validation and new behaviors * Add SIGKILL escalation to plain-node-entry-guard timeout Switch from spawnSync to async spawn to properly handle daemons that trap SIGTERM. spawnSync's timeout only sends the signal and waits, so a daemon that ignores SIGTERM causes the build to hang. The new runDaemonEntry function escalates to SIGKILL after a grace period to enforce the deadline. Configurable timeouts and grace periods via SmokeTimings type; closeBundle hook becomes async to support the change. |
||
|
|
119b1e1c53 |
fix(release): reject non-release publications (#13591)
* fix(release): quarantine unauthorized publications * fix(release): remove unauthorized publication tags * fix(release): require canonical version tags * refactor(release): narrow policy to one workflow * fix(release): reassert latest stable release |
||
|
|
418fbb3192 | fix(mobile-ios): gate release on TestFlight distribution (#13419) | ||
|
|
3b1017c4fb |
Add nightly cut (#13410)
* Add daily macOS dev build release channel Publish once-daily signed macOS builds from main at a dedicated cadence, separate from hourly (too noisy) and release branches (too infrequent). Builds are notarized and installable via the updater, but unvetted — published to stablyai/orca-daily rather than the main repo to avoid evicting stable/RC entries from the releases feed. * fix lint * fix commit * Add third token mint to daily macOS build workflow The upload step's 2x45m retry budget can outlive the one-hour token, so a third is minted after it for verify and cleanup operations. Release notes are moved to a file to ensure consistency between draft creation and publish. Daily channel description updated with specific UTC release time. |
||
|
|
6858e072cf |
fix(terminal): agent pane auto-launch lost under fish + Starship (STA-3417) (#12840)
* fix(terminal): extend the shell-ready startup barrier to fish (STA-3417) Fish never emitted the OSC 777 shell-ready marker, so agent launch commands were written into the PTY while fish/Starship were still initializing: the daemon path wrote them synchronously at session create and the local path blind-wrote ~30ms after the first output byte. The command was echoed by the kernel but never executed. - shell-templates: shared fish --init-command that emits the marker once on the first fish_prompt event (the earliest point fish's own reader owns the PTY, mirroring zsh's zle-line-init marker) - daemon shell-ready: fish joins the startup barrier so the launch command queues until the marker (timeout fallback unchanged) - local-pty-shell-ready: fish launch config gains the marker wrapper - codex-startup-delivery/tui-agent-startup: omp/pi/opencode plans now request shell-ready delivery (codex parity) so the SSH renderer path also waits for the prompt; plain payload-free codex stays on the markerless fast path * fix(terminal): answer DA1 past the shell-ready barrier The barrier queues all inbound input until the ready marker, including the renderer's DA1 reply. A shell that withholds its first prompt until DA1 is answered — fish waits 10s — therefore never emits the marker that would release the reply it is waiting for. Measured: 10.37s to launch an agent, versus 0.35s once the reply lands. Answer DA1 from the daemon while the barrier holds, writing straight to the subprocess so the reply bypasses the queue, and consume the query so the renderer's xterm cannot also reply. Released on ready, timeout, or dispose, handing DA1 back to the renderer for steady state. Consolidates the identical DA1 handler the ConPTY override already used. * fix(terminal): prevent duplicate startup DA1 replies |
||
|
|
850342a3e0 |
fix(ci): run the root-directory guard on stock macOS bash 3.2 (#12879)
* fix(ci): run the root-directory guard on stock macOS bash 3.2 The guard script builds its base-tree lookup with `declare -A`, which needs bash 4+. Its test spawns plain `bash` from PATH, and stock macOS has shipped /bin/bash 3.2 since 2007, so on any Mac without a Homebrew bash the script exits 2 before asserting anything and the default `pnpm test` suite fails 3 of the guard's 4 cases. Machines with a Homebrew bash on PATH never see it, which is why it went unnoticed. Replace the associative array with a plain-array linear scan. Root directories number in the dozens, so the O(n^2) membership check is negligible, and the NUL-delimited reads that protect unusual filenames stay as they were. The empty-array expansion is guarded for `set -u` under bash 3.2. All four guard tests now pass with /bin/bash 3.2; behavior under CI's bash 5 is unchanged. * fix(ci): run the root-directory guard under node instead of bash The guard is the only check in the repo written in shell, and it used `declare -A`, which stock macOS `/bin/bash` 3.2 does not have — so the guard's own test suite failed 3 of 4 cases on any Mac without a Homebrew bash. CI never noticed because runners ship bash 5. Porting it to node removes the interpreter-version variable instead of working around one construct: node is what the sibling script in this directory already uses, it is the runtime that runs the test, and the NUL-delimited read is the same shape as check-changed-code-quality.mjs. It also drops a latent false pass — a failing `git ls-tree` inside the shell's `< <(...)` was not caught by `pipefail`, so the read loop saw nothing and the guard reported success. `execFileSync` throws instead, which is why the two `git rev-parse --verify` probes are no longer needed. Output and exit codes are otherwise unchanged; the usage line now prints node's script path where the shell printed `$0`. Tests pin each guarantee and fail when it is reverted: NUL-delimited reads so odd paths are reported unmangled, exit 2 on bad usage, and git's own 128 with no node stack trace when a sha does not resolve. * fix(ci): keep root entry bytes intact and fence guard output git pathnames are arbitrary bytes, but the guard read ls-tree with encoding 'utf8', so every invalid sequence collapsed to U+FFFD. That mangled the reported name and, because the replacement is not injective, let two different entries compare equal — a genuinely new root entry could be waved through as pre-existing. Read the bytes as latin1 and write them back unchanged. The blocked-entry list is also attacker-controlled and went straight to stdout. The runner trims leading whitespace before matching '::', so an indented entry name still parses as a workflow command, and a pathname may embed a newline. Wrap the list in ::stop-commands:: with a random resume token so only the guard's own annotation is acted on. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
17cfc968cf |
Revert the terminal IME composition-ownership change (#13282)
* Revert "test(ime): restore coverage the composition-ownership change removed (#13168)" This reverts commit |
||
|
|
05c30166f4 |
fix(ci): stop hourly prune from deleting the just-published release (#13266)
* fix(ci): stop hourly prune from deleting the just-published release Hourly prune sorted non-draft releases by createdAt, but nearly every orca-hourly release shares one createdAt from bulk import. At the retain cap, stable sort + reverse put the newest release past the window and immediately deleted it with --cleanup-tag. Sort by publishedAt (tagName as tie-break) and hard-skip the tag this run just published so prune cannot self-delete. * fix(ci): harden hourly prune protect without retain+1 drift Review found that skipping only in the delete loop under-prunes when sort is wrong, and excluding TAG before the retain slice would permanently keep retain+1 releases. Force this run's tag to the front of the sorted list before slicing so it always has a retain seat and oldest builds still prune. Also gate prune on publish_live success and warn if TAG still appears stale. |