Files
tty7/.github/workflows/ci.yml
l0ng-aiandl0ng-ai bed22d899e Keep workspaces whole: remote reopen/restart recovery, and cross-workspace restore guards (#257)
* feat(remote): keep a remote workspace whole across reopens and restarts

Reopening a remote workspace — or coming back to one whose `tty7-server`
had been replaced — landed on a screen of `tty7 — disconnected` panes with
their coding-agent conversations gone. Several independent holes added up
to that; this closes them together, and picks up the surrounding work the
same session produced.

**Telling a restarted server from a blinked link.** `ControlHelloOk` now
carries an `instance` minted once per server *process*. Nothing else in
the handshake changes across a restart — `build` and both dialect numbers
survive it — so a reconnect had no way to know its `pane_id`s were dead.
It does now: a different instance rebuilds the window from its layout
(same tabs and splits, fresh shells in the saved cwds) instead of
re-attaching to a process that is gone. An absent instance means *unknown*
and is never read as a restart.

**An attach can now fail.** `Attach` has no synchronous reply, so the
client returned `Ok` unconditionally and the daemon's `Error` frame was
read much later by the reader thread, which has no arm for it — the pane
then landed in the *link is down* state instead of falling back to a fresh
shell. The client now reads far enough into the reply to classify it on
the kind byte (the snapshot behind it can be megabytes) and hands those
bytes to the reader thread, so a successful attach loses none of its
replay. Local and remote attaches get different waits: the local one is on
the UI thread.

**The agent session survives to be resumed.** `TerminalView` raises
`AgentSessionChanged` when the pane's agent reports a new native session
id, so the layout on file catches up instead of waiting for the user to
happen to open a tab. A pane that is still connecting now carries its
agent through `PendingSpawn` — a save landing in that window used to write
`agent: null` over the record — and `land_pane` sends `--resume` when the
attach turned out to need a fresh shell.

**Ending sessions says so on file.** "End Sessions" kills the panes and
then drops their ids from the record, pushing the cleared layout to the
machine that owns it (design §10: the remote's copy wins, so a local-only
clear would be undone by the next open — the open this exists for).

**The new-tab dropdown lists the window's machine.** `Host::shells` and a
`Shells` control request (dialect v2) make the "+" menu a property of the
machine the window is bound to. A remote window filled from this
computer's `/etc/shells` offered `/bin/zsh` on a box whose zsh is
elsewhere, and every pick failed to spawn.

**An install reports its bytes.** The download and the SFTP upload each
report progress, relayed to the client over the routed connection as a
`RoutePrompt::InstallProgress`, and painted as a bar under the machine's
row in the switcher. ~8 MB across two hops behind the word "connecting…"
was indistinguishable from a hang.

**The installer compares dialects, not version strings.** `tty7-server
--protocol` prints what a binary speaks without starting it, so a connect
adopts an already-running server it can talk to rather than prompting
about a build difference and uploading 8 MB the machine did not need.

**Switcher.** A machine's `⋯` menu holds "New Workspace" (it was a row
under every machine, pushing the list a quarter of a card down) and a new
"Disconnect", which drops the connection and leaves the windows open and
read-only. The suspension lasts exactly as long as that machine has a
window on it.

Also drops three design/contract docs for the now-shipped remote-workspace
work.

* fix(session): stop one workspace's panes from being restored into another

A restart put a copy of one workspace's seven tabs — cwds, layout and
recorded agent sessions — in front of another workspace's own tabs, and
auto-resumed every one of those agents a second time: six `claude
--resume <id>` pairs running in parallel against the same conversations,
one set per window. The record-level corruption that seeded it is still
unattributed, but every mechanism that let it propagate, amplify, or go
unnoticed is closable, and this closes them.

**Panes now know their owner.** `Spawn` can carry the workspace the pane
is created for; the daemon stores it immutably and reports it in
`List`'s `PaneInfo.owner`. Restore refuses to re-attach a pane another
workspace owns (`pane_attachable`) — before this, a saved id landing on
somebody else's live pane attached silently, which is how one window
could pick up another's shells. The field rides a new `SPAWN_OWNED`
frame with a struct payload (the legacy spawn payloads are positional
tuples an old daemon cannot grow), gated on a new `pane-owner` feature
string: a client only sends it to a daemon that advertises it, so the
legacy kinds stay byte-for-byte what old daemons expect. A pane with no
recorded owner stays attachable by anyone — that is the pre-field
behavior, not a new risk.

**Saved pane ids are bound to the daemon process that issued them.**
`DaemonVersion` now carries an `instance` minted once per process (the
local twin of the control hello's), the GUI caches it at the
`ensure_running` handshake, and each local workspace records it as
`daemon_instance` beside its layout. Claiming a workspace whose ids came
from a different instance blanks them first: daemon pane ids restart
from 1, so after a reboot every saved id points at whatever unrelated
shell holds the number now, and the aliveness check cannot tell a
survivor from a squatter. A blank on either side means "cannot tell" and
never trips it. Unlike the duplicate-claim case below, this path keeps
the agent resume — the pane is genuinely gone with its daemon, and the
fresh shell resuming the conversation is the feature.

**A duplicate claim loses its agent resume along with its pane id.**
`dedupe_pane_ids` kept the loser's layout *and* its
`agent_session_id`, so the blanked leaves took restore's spawn-fresh
path and auto-typed `claude --resume` for conversations the winning
workspace's panes were still running — the doubling above. The winner
keeps the panes and the resume; the loser keeps only cwds.

**Cross-workspace saves are caught at the write.** Every terminal view
remembers the workspace whose window created it, and `save_session`
logs an error naming both ids if a window ever records a pane created
for a different workspace — the tripwire for the still-unattributed
seed corruption, so a recurrence is caught in the act instead of
reconstructed from `session.json` archaeology days later.

Wire compatibility both ways: `PaneInfo.owner`, `DaemonVersion.instance`
and `Workspace.daemon_instance` are `#[serde(default)]` struct fields
(old peers' JSON decodes, new fields are ignored by old readers), and
`SPAWN_OWNED` is feature-gated as above. `daemon_instance` is
client-owned in the design-§10 storage split — it names the local
daemon, and the field-census test pins the classification.

* fix(session): resume the agent when a local pane dies mid-restore

`session_to_pane` decided whether to send a coding agent's `--resume`
from `restore.is_none()` — i.e. from whether the pane looked alive when
the restore started. But `alive_panes_on` runs one `List` at the top of
the restore, while the attaches happen per leaf afterwards. A pane that
exited in between failed its attach, fell back to a fresh shell inside
`spawn_shell_terminal_in`, and then landed in the `restore.is_some()`
arm: an empty shell with its conversation dropped.

`ShellParts.restored` already answers this exactly, and the remote path
already reads it in `land_pane`. Carry it onto `TerminalView` so the
synchronous local path can read it too, and branch on that instead of
re-deriving the answer from a set that may be stale by the time it is
used.

No behaviour change on the paths that were already correct: a view that
was never restoring anything reports `restored: false`, which is the
same answer `restore.is_none()` gave them.

* fix(remote): check the server instance against the record, not just memory

A remote workspace's pane ids were only guarded against server restarts
by `RemoteLinks::instances`, an in-memory map. On the first connect after
the client starts, every machine is a first sighting, so `server_restarted`
answers false — and a `tty7-server` that was replaced while the client was
closed sails straight through. Its pane ids restart from 1, so the saved
ones now name unrelated shells, and the reconnect attaches to them: the
exact id-reuse failure the local side already guards against.

`Workspace::daemon_instance` was local-only for the stated reason that a
remote server's identity is tracked live per connection. That tracking is
correct but not sufficient — it cannot survive the client restart that
makes the question worth asking.

So the field now means the same thing on both sides: which process minted
the pane ids in this record. `WorkspaceStore::serving_instance` picks the
local daemon or the far machine's server depending on the workspace, and
`finish_attempt` compares it per workspace before deciding to re-attach or
rebuild. It stays client-owned: it records what *this* client last saw, so
two clients on one remote workspace each keep their own and neither may
overwrite the other's.

An unreachable machine still records nothing, which is what keeps a good
stamp from being erased with `None` — that would disarm the next check.

Also in these three files: the §N references to the deleted design docs,
cleaned up as part of the sweep in the following commit.

* docs: drop the references to the deleted design documents

The three documents this branch removed were cited ~280 times: `design
§10`, `contract §8`, `§17` and friends in comments, five references by
file path in code and manifests, five in CI workflows and one in the
release skill. Every one of them now points at nothing.

Rewritten rather than merely stripped, because most were not decoration:
"design §10 makes the remote's `workspaces.json` the authority" becomes a
statement in its own right, and the several that carried a Chinese phrase
from the document as their justification say the same thing in English
instead. Where the reference was purely parenthetical it is simply gone.

Not touched: `PRD §7.1`, `brief §8` and the like, which name documents
this branch did not remove and were already external before it, and the
`RFC 4648 §10` test-vector citation, which is a real specification.

The `host boundary` CI job loses `(§10.6)` from its name. It is not one of
the required checks, so branch protection is unaffected.

---------

Co-authored-by: l0ng-ai <24760907+l0ng-ai@users.noreply.github.com>
2026-07-29 19:15:19 +08:00

185 lines
7.8 KiB
YAML

name: CI
# Compile + test on every push/PR. The Windows and Linux jobs are the
# compile-feedback loop for the platform-specific code a macOS dev machine
# never builds (`cfg(windows)` transport / process detach / config dir,
# the Linux `/proc` queries, the x11/wayland gpui backends). The macOS job
# guards against regressing the original target.
on:
push:
branches: [main]
pull_request:
workflow_dispatch:
jobs:
fmt:
name: rustfmt
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
with:
components: rustfmt
- run: cargo fmt --check
# `ui::` and `terminal::` must not reach the filesystem or git
# directly, because once a workspace can be remote those calls answer for the
# wrong machine. The allowlist of genuinely-local paths lives in the
# script, next to the reason each one is exempt.
#
# A standalone job on purpose, and one that must stay *non-required*: main's
# required checks are `rustfmt` and the three `build & test (<target>)` names,
# and folding this into either would make it required the moment it lands —
# wedging every open PR on a check they have never seen. Same reasoning as
# `server-musl` below. Cheap enough (a checkout and a grep) that it does not
# need caching or a toolchain.
host-boundary:
name: host boundary
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: bash .github/scripts/check-host-boundary.sh
build:
name: build & test (${{ matrix.target }})
strategy:
fail-fast: false
matrix:
include:
- runner: macos-14
target: aarch64-apple-darwin
- runner: windows-latest
target: x86_64-pc-windows-msvc
- runner: ubuntu-latest
target: x86_64-unknown-linux-gnu
runs-on: ${{ matrix.runner }}
steps:
- name: Checkout tty7
uses: actions/checkout@v4
# gpui-component is a git dependency (see Cargo.toml), so no sibling
# checkout is needed. The Windows backend (gpui_windows + DirectWrite/D3D)
# ships with the windows-latest runner's SDK — no extra system deps.
# gpui's Linux backends need the x11/wayland/xkb/font dev packages at
# build time (build scripts resolve them via pkg-config). Same set the
# README documents for building from source on Linux.
- name: Install Linux system dependencies
if: runner.os == 'Linux'
run: |
sudo apt-get update
sudo apt-get install -y pkg-config cmake clang libxkbcommon-dev \
libxkbcommon-x11-dev libfontconfig1-dev libfreetype6-dev \
libwayland-dev libx11-dev libxcb1-dev libzstd-dev libssl-dev \
libkrb5-dev
echo "LIBGSSAPI_IMPL=mit" >> "$GITHUB_ENV"
- uses: dtolnay/rust-toolchain@stable
with:
targets: ${{ matrix.target }}
- uses: Swatinem/rust-cache@v2
# `--locked` on the build so a Cargo.lock that disagrees with Cargo.toml
# fails here instead of being silently rewritten. Without it the drift is
# invisible to CI and lands on contributors instead: every local `cargo`
# run rewrites the lock, leaving a permanently dirty working tree that has
# to be re-discarded before every commit. The release workflow locks its
# build too; only nightly stays unlocked, because it stamps Cargo.toml's
# version and relies on cargo refreshing the lock's own root entry.
- name: Build
run: cargo build --locked --target ${{ matrix.target }}
- name: Test
run: cargo test --locked --target ${{ matrix.target }}
# Static musl builds of the headless server binary that remote workspaces push
# onto the far machine (decision D10).
# One binary has to run on any distro without regard to the target's glibc
# version, so it is linked fully static against musl rather than built per
# distro. Compile-only — this job publishes nothing; release.yml and
# nightly.yml carry the same job with an upload step.
#
# Deliberately a *separate* job rather than two more rows in the `build` matrix
# above: those three `build & test (<target>)` names are main's required
# checks, and reshaping that matrix would wedge branch protection on every open
# PR. Keep this job non-required until it has a few weeks of green.
server-musl:
name: tty7-server musl (${{ matrix.target }})
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
target:
- x86_64-unknown-linux-musl
- aarch64-unknown-linux-musl
# `-C strip=symbols` is applied by rustc through the linker, so it strips the
# aarch64 output from an x86_64 runner — plain `strip(1)` would not. Set here
# rather than in [profile.release] because the release profile is shared with
# the GUI builds, which keep their symbols.
env:
RUSTFLAGS: -C strip=symbols
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
with:
targets: ${{ matrix.target }}
# zig supplies both the musl sysroot and the C cross-compiler, which is
# what makes one x86_64 runner able to emit both musl targets. russh's
# default crypto backend (aws-lc-rs) builds a sizable C/asm library through
# cmake, and that is the part every other approach trips over: `cross`
# needs a custom image to get cmake into the sandbox, and Ubuntu's
# `musl-tools` only ships an x86_64 `musl-gcc` wrapper with no C++ driver
# and nothing at all for aarch64. Verified locally: both targets link
# statically and the resulting binaries run under Alpine.
#
# If setup-zig ever becomes a problem (its GitHub repo is a mirror of a
# Codeberg original), cargo-zigbuild also accepts zig from PyPI —
# `pip3 install ziglang`, which it finds via CARGO_ZIGBUILD_PYTHON_PATH —
# so this can drop to zero third-party actions without changing anything
# else.
- uses: mlugg/setup-zig@v2
with:
version: 0.16.0
- uses: taiki-e/install-action@v2
with:
tool: cargo-zigbuild
- uses: Swatinem/rust-cache@v2
with:
key: ${{ matrix.target }}
# The crate split lands separately; until `tty7-server` exists as a
# workspace member this job has nothing to build. Skip cleanly rather than
# fail, so the workflow can land before the split and simply start working
# once it arrives. `--no-deps` keeps this to a manifest parse — no
# resolution, no network.
- name: Look for the tty7-server package
id: probe
run: |
set -euo pipefail
if cargo metadata --no-deps --format-version 1 \
| jq -e '[.packages[].name] | index("tty7-server")' >/dev/null; then
echo "present=true" >> "$GITHUB_OUTPUT"
else
echo "present=false" >> "$GITHUB_OUTPUT"
echo "::notice::tty7-server is not a workspace member yet (crate split, M1) — nothing to build"
fi
# `-p tty7-server` addresses the package by name, so this survives whatever
# directory layout the split settles on. It also keeps feature unification
# scoped to the server's own dependency graph: the GUI crate is the one
# that turns on tty7-core's `gssapi` feature, and that feature cannot build
# under musl (libgssapi-sys wants a system MIT/Heimdal krb5). Building the
# whole workspace here would drag it in and fail.
- name: Build static tty7-server
if: steps.probe.outputs.present == 'true'
run: cargo zigbuild --release --locked -p tty7-server --target ${{ matrix.target }}
- name: Assert the binary is static
if: steps.probe.outputs.present == 'true'
run: bash .github/scripts/assert-static.sh "target/${{ matrix.target }}/release/tty7-server"