Commit Graph
9708 Commits
Author SHA1 Message Date
Neil 20b3544328 test(linux): record native AppImage startup evidence 2026-09-01 03:26:36 -07:00
Neil 486d04e246 test(linux): reject non-executable AppImages 2026-09-01 03:26:35 -07:00
Neil 0c86c83c96 fix(linux): keep local package targets aligned 2026-09-01 03:26:34 -07:00
Neil 51227b6182 fix(linux): make local package architecture explicit 2026-09-01 03:26:33 -07:00
Neil 462978fba6 fix(linux): bind AppImage names to architectures 2026-09-01 03:26:32 -07:00
Neil aed01aa935 fix(linux): harden AppImage desktop startup oracle 2026-09-01 03:26:30 -07:00
Neil 9a3260197b test(linux): preserve AppImage startup failure logs 2026-09-01 03:26:29 -07:00
Neil 6655fb6148 fix(linux): bind AppImage runtime and desktop startup checks 2026-09-01 03:26:28 -07:00
Neil 4cffd9a8f9 fix(ci): route headless pairing package contracts 2026-09-01 03:26:26 -07:00
Neil 299f7296f2 fix(linux): verify static AppImages before publishing 2026-09-01 03:26:25 -07:00
Neil d92288d11b fix(linux): require static AppImage runtimes 2026-09-01 03:26:23 -07:00
Neil e241a3fb10 test(linux): assert on CLI output, not the harness's own control lines
run-cli-case.sh echoes `RESULT status=N case=<name>`, and the two cases named
*-skills asserted `expectOutput: 'skills'`. That substring was satisfied by
the case name in the harness's own line, so 2 of 8 cases asserted nothing
about the command -- gutting `skills` entirely would still have gone green.

Control lines are now excluded before matching, and both cases assert the
rendered help header, which only real help output produces. Verified on an
Ubuntu 24.04 host: 8/8 still pass against a stack-tip AppImage.

Also register the gate in reliability-gates.jsonc, which #15085 added a CI
Docker gate without. Red/green is recorded from a stock release AppImage
failing 4 of 8, three of them at status 133 (SIGTRAP).
2026-09-01 03:26:20 -07:00
Neil e3b526056f test(linux): bound shutdown cleanup grace 2026-09-01 03:26:19 -07:00
Neil 0aaa0be45f test(linux): poll shutdown readiness without tail leaks 2026-09-01 03:26:17 -07:00
Neil b4867766d8 fix(ci): route all Linux packaging contract changes 2026-09-01 03:26:15 -07:00
Neil ee6e509f21 ci(linux): give package contracts timeout headroom 2026-09-01 03:26:13 -07:00
Neil dc56e2b038 test(linux): add startup margin to shutdown oracle 2026-09-01 03:26:11 -07:00
Neil a190aa1b93 test(linux): tolerate readiness timeout boundary 2026-09-01 03:26:09 -07:00
Neil d1e5013f99 test(linux): signal AppImage serve owner directly 2026-09-01 03:26:08 -07:00
Neil e5800e63b7 test(linux): avoid buffered serve readiness detection 2026-09-01 03:26:07 -07:00
Neil 9b6b145e4a test(linux): add a packaged-artifact contract for the CLI launch paths 2026-09-01 03:26:06 -07:00
Neil ee41acebc5 fix(linux): recognise abstract X sockets and inherited Wayland fds
Two display setups this gate could not prove were refused outright, and on the
desktop path that is app.exit(1) with no workaround.

An X server may bind only the abstract namespace (`@/tmp/.X11-unix/X0`), which
leaves no filesystem socket to stat. Abstract addresses are kernel-owned and
vanish the moment the owner exits, so an entry in /proc/net/unix is proof of a
live server -- no lock file needed and no stale entry possible. Verified on
Ubuntu 24.04, where 139 such addresses were present.

WAYLAND_SOCKET is an already-connected fd handed over by the compositor, so
there is no path to stat and WAYLAND_DISPLAY may be unset entirely. Its
presence is the display.

Both are consulted only after the filesystem-socket check fails, so no
existing verdict changes.
2026-09-01 03:26:04 -07:00
Neil 69c13fc30b fix(linux): fail closed when a stale socket blocks the Xvfb rebind
Readiness only checked that /tmp/.X11-unix/X99 exists. A stale socket we
could not unlink still exists after our own Xvfb refused to bind, so Orca set
DISPLAY to a dead server and Chromium died in Ozone init.

Measured on Ubuntu 24.04 against the pre-fix build: with a leftover :99
socket and no lock, serve exits 139 (SIGSEGV), the socket inode is unchanged
before and after, and no lock is recreated -- it neither cleaned up nor
respawned. To a user that is a crash, not a misconfiguration.

This is reachable in the documented topology, where orca-xvfb.service has no
User= and runs as root while serve runs as User=orca: /tmp is sticky, so the
orca uid cannot unlink a root-owned socket, rmSync fails, and Xvfb exits with
the display already active.

Readiness now requires the display to actually be live -- our socket plus a
lock naming a running process -- so the same state reports an unusable
display and exits 1 with the existing diagnosis.
2026-09-01 03:26:03 -07:00
Neil c0bc64fc86 fix(linux): do not treat a lockless X socket as a dead display
An X server writes its lock beside its socket and both survive a crash
(verified against Xvfb under SIGKILL), so a socket with no lock was never
left by a crashed server. It is an endpoint published from elsewhere: a
container bind-mounting only /tmp/.X11-unix, WSLg, or a foreign PID
namespace. Declaring those dead made the desktop gate exit(1) on displays
that work, with no workaround, and the serve gate refuse to start.

Liveness now splits by ownership. A foreign DISPLAY trusts a lockless
socket; Orca's own :99 does not, because removeStaleDisplayArtifacts
unlinks the lock before the socket and so manufactures that state itself --
adopting it would resurrect the orphan-socket bug and stop the cleanup from
self-healing. The stale-lock rejection is unchanged.

Also correct four doc statements this behaviour falsified.
2026-09-01 03:26:01 -07:00
Neil 48de327860 fix(linux): fail serve when no display is available 2026-09-01 03:26:00 -07:00
Neil d6ad56e036 test(packaging): split runtime resource checks 2026-09-01 03:25:58 -07:00
Neil 0388969ee0 chore: format reliability gate manifest 2026-09-01 03:25:57 -07:00
Neil 3695bfc4cb fix(linux): preserve unverified external displays 2026-09-01 03:25:56 -07:00
Neil 4437674a2f refactor(linux): read display locks without a preflight race 2026-09-01 03:25:55 -07:00
Neil 8d108d3fe8 fix(linux): report a missing display instead of dying in uv_close 2026-09-01 03:25:53 -07:00
Neil 4f90ecf001 fix(linux): strip injected Chromium switches from CLI args 2026-09-01 03:25:50 -07:00
Neil e8c336baaf fix(linux): respect CLI flag value boundaries 2026-09-01 03:25:49 -07:00
Neil d5f6f7e68e fix(linux): tighten CLI launch detection 2026-09-01 03:25:48 -07:00
Neil cd298ce165 test(linux): cover AUR serve wrapper flags 2026-09-01 03:25:48 -07:00
Neil 56977eb04f fix(cli): redirect the open-url command before startup 2026-09-01 03:25:47 -07:00
Neil 8488259dca test(cli): cover command-named project selectors 2026-09-01 03:25:46 -07:00
Neil 53639d6c81 refactor(cli): remove redundant command membership check 2026-09-01 03:25:45 -07:00
Neil 0df64c375d fix(linux): stop CLI commands from falling through to Chromium startup 2026-09-01 03:25:44 -07:00
Neil 312cf01bc7 fix(linux): stop re-extracting the AppImage on inode metadata churn
The extracted-payload cache key hashed ctime alongside dev/ino/size/mtime.
ctime moves on any inode metadata write -- `chmod +x`, which every AppImage
user is told to run, plus `chown`, an ACL or SELinux relabel, and a backup
restore -- none of which alter a byte of the payload.

Measured on Ubuntu 24.04: `chmod +x` leaves dev, ino, size and mtime
identical and moves ctime alone, so the key changed and the next launch paid
a full ~519 MB re-extraction and a multi-second stall to rebuild a payload it
already had, then pruned the old generation.

Key on content identity instead. An in-place content change moves mtime and
almost always size; a replacement moves the inode. The existing
replace-in-place test still passes.
2026-09-01 03:24:51 -07:00
Neil 1fb705064c fix(linux): bound the CLI registration lock wait
`retries: 1000` caps the attempt count, not elapsed time, so at up to 1s per
attempt an IPC-driven registration could hang ~16 minutes against a wedged
holder with no feedback.

A legitimate holder is bounded by the extraction timeout, so wait that plus
slack and then fail with a message naming the lock file, rather than hanging.
`maxRetryTime` is forwarded verbatim to the `retry` package by proper-lockfile.
2026-09-01 03:02:30 -07:00
Neil c23953a5a0 fix(linux): reclaim superseded AppImage payloads and packaged symlinks
Pruning removed 3215 of 3216 files from a superseded generation and always
stranded resources/app.asar, leaking ~105 MB per version update. Electron's
asar shim reports a *.asar file as a directory, so the recursive remove tried
to rmdir a real file and failed with ENOTEMPTY; the .catch(() => {}) hid it.
Reproduced end to end on Ubuntu 24.04: 519M -> 623M across one update, and
519M again once the payload is actually reclaimed.

removeExtractedAppImagePayload holds process.noAsar for the removal, counted
so overlapping removals cannot hand the shim back early, and the prune site
now warns with the path instead of swallowing the rejection. All three
removal sites use it -- staging cleanup and displaced roots leaked the same
way.

Also reclaim symlinks left by a packaged deb/rpm install, which the
extracted-cache-only rule turned into a hard conflict on a deb -> AppImage
migration, and name the remedy in the conflict error.
2026-09-01 00:47:52 -07:00
Neil 15e13881da refactor(linux): import bundled launcher directly 2026-08-31 05:48:30 -07:00
Neil cc575461c9 docs(linux): make headless AppImage extraction runnable 2026-08-31 05:48:29 -07:00
Neil f821ca0cca fix(linux): accept extracted AppImage runtimes with APPDIR only 2026-08-31 05:48:29 -07:00
Neil daff0b9301 fix(linux): fence AppImage terminal shim mounts 2026-08-31 05:48:29 -07:00
Neil 6c1a6d39c1 test(cli): assert registration lock serialization 2026-08-31 05:48:29 -07:00
Neil 657967f30a refactor(linux): trim AppImage CLI registration seams 2026-08-31 05:48:29 -07:00
Neil 307e9f0f7d fix(linux): give the CLI one entrypoint by extracting the AppImage once 2026-08-31 05:48:28 -07:00
github-actions[bot] 212c0e42a6 Update README downloads badge 2026-08-31 12:37:51 +00:00
Neil 75e5c996c1 perf(relay): stop ACK boundary scans at first pending boundary (#17491)
* perf(relay): stop ACK boundary scans at first pending boundary

* test(relay): pin PTY source boundary cleanup and guard ascending sends

The early-`break` in advanceCredit is only correct while sentBoundaries is
inserted in ascending sentEndSu order. Turn that implicit invariant into a
throw at the sole live write site (commitPtySourceSend), and assert the
post-state directly instead of inferring it from an iteration budget:

- assert the surviving boundary set after the 1,023-ACK benchmark
- cover the jump-ahead cumulative ACK that must delete many boundaries in
  one pass (the case an over-eager `break` would get wrong)
- cover the settleReservedPtySourceAck -> advanceCredit entry point
- drop an arithmetically-implied assertion and CI benchmark log noise

* perf(relay): reclaim ACK boundaries with a monotone cursor

The early-break Set scan still rebuilt a Set iterator per ACK, so V8 walked
delete tombstones and the drain stayed superlinear; the visit-count test could
not see it because it stubbed sentBoundaries with a generator over a private
Set. Replace the Set with an ascending boundary list plus a monotone cursor,
assert the real structure, and add a benchmark over the shipped code.

* test(relay): enforce ascending sent-boundary inserts in the collection

Move the ascending-order precondition into PtySourceSentBoundaries.add so
both insert sites are covered, and assert per-ACK span reclamation in the drain.

* test(relay): collapse ledger test record accessors into getDeliveryRecord

Rebase onto #17490 left two structurally identical internals accessors
(getCursorRecord, getBoundaryRecord); one typed accessor covers both.
2026-08-31 03:10:44 -07:00