29 Commits
Author SHA1 Message Date
ayamirandl0ng-ai e231b16fb3 feat(ssh): allow remote image clipboard writes (#766)
* feat(ssh): allow remote image clipboard writes

* fix(ssh): keep a profile's clipboard grant across a re-attach

A native ssh pane's OSC 5522 permission is decided by the spec that
dialled the host, and the daemon is the only side that holds it. A window
reopening onto a pane that outlived it attaches by pane id, has no spec
to read, and sends `allow_remote_clipboard_write: false` — which the
daemon took as the new answer and the pane's own view took as a refusal.
Both sides then said no, so the first restart after switching the
permission on turned every copy into an `EPERM` with the switch still
reading "on".

Pin the spec's answer in the pane and route both attach and detach
through one decision point, so a pane that carries a spec keeps that
spec's answer whatever an attaching client claims, and a pane without one
— everything on a remote `tty7-server` — is exactly as permitted as its
controller says. On the client side, refuse only what the pane can see is
forbidden and leave the verdict to the daemon otherwise.

Also: release a failed transfer's buffered bytes instead of parking up to
`MAX_CLIPBOARD_BYTES` per pane until the next request, and answer the
capability probe with the permission actually in force rather than a
constant that always reads as "off".

---------

Co-authored-by: l0ng-ai <24760907+l0ng-ai@users.noreply.github.com>
2026-09-02 14:12:28 +08:00
l0ng-ai 3dc63e2d87 fix(daemon): find and reap a seat-holding daemon that lost both its names (#671)
A daemon can survive quit-and-stop with its endpoint unlinked and its
pidfile gone while still holding the singleton seat (#667). Every later
launch then spawns a daemon that stands down against the lock and times
out red, and nothing on the machine can recover: stop answers "not
running", ensure_running reaps only through the pidfile, and flock
cannot say who the holder is.

Two roads led there, and both are closed:

- The reap identified a daemon by proc_pidpath alone, which fails
  outright for a live process whose binary was deleted — every nightly
  update replacing the installation. The identity check now falls back
  to the kernel's comm name (proc_name on macOS, /proc/pid/comm on
  Linux, both recorded at exec and immune to deletion), strips Linux's
  " (deleted)" marker, and — decisively — no longer deletes the
  pidfile of a live process it cannot identify: the record was the only
  handle left on the survivor.

- When the pidfile is gone entirely, the pid the claimant now writes
  into daemon.lock at claim time is the handle of last resort. The lock
  file is never deleted and holding the flock is the definition of
  being the server, so while the seat is held its content names the
  holder; stop() and the reap fall back to it, and a confirmed reap
  clears the record (only under a momentarily-free seat) so a stale
  number cannot outlive its process. Unix-only: the Windows seat is
  share_mode(0), unreadable while held.

Every road back now clears a stranded seat, not just the GUI's:
ensure_running's stale cleanup is factored into spawn::reap_stranded,
tty7 server start runs it too, and tty7 server stop no longer takes
"nobody answered" for "nothing to stop" when the seat is still held.
A short grace keeps the reap away from a daemon that is merely
mid-handoff or mid-startup — where health is an answered handshake,
never a bare connect: a wedged daemon's listener still completes
connections out of the kernel's backlog. The startup-timeout errors
name the seat-holding pid, with the kill advice identity-gated so a
stale record never tells anyone to kill an innocent process.

Two liveness corrections round it out: a zombie now reads as dead — it
answers kill(pid, 0) like the living but holds no lock and no image,
and no signal can end it, so counting it alive spent both reap timeouts
on a corpse (the GUI never waits on the daemons it spawns, so crashed
daemons are zombies as a rule) — and stop() only pays the
process-exit wait for a shutdown it actually delivered, instead of
watching an unreached survivor not move for five seconds.

The guard tests were each verified to fail against the behavior they
pin (fallbacks, the handshake criterion, the grace, and the wait gate
removed by mutation) before being trusted green; the zombie probe
semantics (proc_pidinfo failing for a zombie that still answers signal
0) were measured, not assumed.
2026-08-17 15:11:16 +08:00
l0ng-aiandl0ng-ai 0df604054d fix(daemon): keep a lingering daemon findable and reapable after quit-and-stop (#655)
* fix(daemon): keep a lingering daemon findable and reapable after quit-and-stop

Quit-and-stop could strand a daemon that had already unlinked daemon.sock
and deleted daemon.pid but never finished exiting: libc exit() runs atexit
handlers and static destructors beside dozens of live threads, and a
finalizer that blocks leaves the process holding the singleton lock with no
name on disk. Every later launch then spawns a daemon that stands down
against the lock and times out red, forever.

Three changes, each a fallback for the others:

- on_shutdown keeps the pidfile: once the endpoint is unlinked it is the
  only handle anything has on a process that is not gone yet. A pidfile
  that outlives a clean exit was already handled by recorded_daemon_is_dead
  and the reap path.
- The daemon exits through _exit(2) (after flushing the logger), skipping
  the atexit/destructor window entirely; everything owed to disk is flushed
  explicitly in on_shutdown.
- spawn::stop reaps with the pid it captured before asking the daemon to
  die, instead of re-reading a pidfile an old build's shutdown may have
  wiped mid-stop; reap_recorded_daemon keeps the pidfile when the process
  survives even SIGKILL, so the next attempt still has someone to reap.

* review: fix stale stop() comment, pin the mid-stop pidfile-vanish ordering in the test

The comment at the top of stop() still claimed a clean shutdown removes
the pidfile, which this branch just made untrue; it now states the real
reasons the pid is captured early. The vanishing-pidfile test now asserts
the sweeper's delete actually landed while stop() was waiting, so a
future shrink of PROCESS_EXIT_TIMEOUT cannot silently turn it into a
weaker scenario.

---------

Co-authored-by: l0ng-ai <24760907+l0ng-ai@users.noreply.github.com>
2026-08-16 17:35:39 +08:00
webdev 84a424006e fix(resize): defer the reflow to the daemon's Size echo on remote routes too (#632)
Since #415 the daemon echoes a `Size` frame at the stream position where the pty changes geometry and the client defers its grid reflow to that marker — but only on local routes, so a remote pane resized mid-flood still parsed queued old-width bytes into the new-width grid, which network transport makes worse.

Rather than probing `Version` per pane (a whole routed connection, and for ssh/WSL a whole bridge process, on every spawn and attach), the server advertises the pane protocol's features on its control hello. The answer is cached on the link and read off the host when a pane's route is built, and the route carries it to the terminal at spawn, attach and relink. This is additive within `CONTROL_VERSION` 7: no new field, just extra names in the existing `ControlHelloOk.features`, so an older client cannot choke and an older server that names no echo makes the client reflow at request time as before. A route built while the link is down answers false.

Known limitation, inherited from #415's design and not introduced here: there is no timeout if a promised echo never arrives — once deferred, a later identical resize neither re-sends nor reflows, so a wrongly-set bit would freeze the grid at the old geometry. Every traced path makes the control hello and the pane daemon the same build, normally the same process.

Closes #416.
2026-08-14 16:02:04 +08:00
l0ng-ai a2eab07f8d fix(ssh-config): watch every included file an alias was parsed from (#528)
The cache behind `alias_still_resolves` keyed on the mtime of
`~/.ssh/config` alone, and `Include` leaves the including file's timestamp
untouched — so an alias put back in the file it lived in went on reading
as gone, parking its workspaces with no retry and no error. Every file the
parse reads is watched now, plus the root even when unreadable.

Two tests that failed on CI for reasons outside the code they cover go
with it. `pane_history` waited on the seed file existing, but the snippet
seeds through a redirection, which creates the file before `tail` fills
it. `daemon::singleton` asked for the seat back the instant it dropped it,
and a neighbour's fork keeps an inherited descriptor referencing the same
flock until its exec.

Closes #524. Closes #525.
2026-08-11 22:53:29 +08:00
l0ng-aiandl0ng-ai 425f87e9a4 fix(core): key the machine tree to the config directory (#462)
* fix(core): key the machine tree to the config directory

The tree resolved from $HOME while everything else an instance owns —
views.json, the scrollback, the history, both sockets, the pidfile, and
daemon.lock — resolved from the config directory. So --config-dir moved
every part of an instance except the one that says which workspaces
exist, and two tty7s pointed at different config directories, each
holding its own lock and each certain it was the only server on the
machine, still co-owned one ~/.local/share/tty7/machine.json.

MachineStore::persist writes the document whole. The second one to flush
replaced the first one's workspaces with its own, and the next daemon to
start read the survivor's tree as the machine's. An empty tree is not
distinguishable from a machine that really has nothing on it, so the GUI
does what an empty tree means and forgets those workspaces for good.

A lock and the thing it protects have to be keyed alike. data_dir() now
follows the config directory; TTY7_DATA_DIR stays as the highest-priority
override so the test harnesses keep their sandboxes.

Moving the path without carrying the file would lose every workspace at
the moment of upgrade, which is the failure this change exists to stop,
so the daemon adopts the legacy file on startup before it opens the
store. The destination already existing is the whole guard: it means a
newer run owns the tree and the copy at the old path is stale, from a
build that predates the move and still writes where it believes the tree
lives. Adopting that over the live file would hand the old tree back.

* fix(core): only the machine's own instance inherits the legacy tree

The migration moved `machine.json` into whichever config directory started
first. In the very setup this change exists to fix — a default install beside a
`--config-dir` one — that is the second instance renaming the machine's tree
into its own directory, leaving the primary to come up owning nothing. It also
fired in our own test suite, where `routed_pane` and friends launch a real
`tty7-server --config-dir <TempDir>` under the developer's own `HOME`.

Adoption is now the entitlement of the instance running out of the config
directory this machine resolves to on its own: `$TTY7_CONFIG_DIR` where the box
names one, `$HOME`'s otherwise. Comparing paths rather than asking whether
`--config-dir` was passed is what keeps the ordinary install working — `spawn`
hands every daemon it starts an explicit `--config-dir`, its own included — and
counting `$TTY7_CONFIG_DIR` is what keeps remote hosts upgrading, since a remote
`tty7-server` is launched without the flag and finds its directory that way.

Also tightens the cross-filesystem fallback: a rename that failed because
another process already carried the file over is the one benign race, not an
error to report and not something to copy over. What is left copies through
`create_new`, so "never overwrite what is already there" holds against a racing
writer and not merely against an `exists` check several syscalls old, and a
write that does not finish leaves nothing behind.

Tests: the gate both ways, the appearance hint riding along, the same directory
under two names, the copy path refusing an occupied destination, and two
cross-process cases in `machine_tree` that start a real server under a scratch
`HOME` — one carrying the legacy tree in, one leaving it alone.

---------

Co-authored-by: l0ng-ai <24760907+l0ng-ai@users.noreply.github.com>
2026-08-10 17:12:26 +08:00
l0ng-ai a55340ed7f fix(daemon): sweep a dead daemon's leavings on the writer's tick, not at startup
Review follow-ups on this branch.

`history::sweep` still ran at startup, three lines under a new comment
explaining why sweeping there is wrong. The reasoning transfers exactly, and
worse than by analogy: a restore carries the dead pane's commands to its
successor via `history::carry`, so sweeping before the window can ask deletes
the file the request is about. Same shape as the scrollback bug, one file over.
Both sweeps now run on the writer's tick off one shared id set, and the writer
is named for what it does.

`pane_attachable` lost its only caller when the restore path moved to
`pane_free_for`, leaving a function kept alive by the test asserting on it. The
attach site does not need to predict the listing: it tries the attach, and a
pane that is gone falls through to the fresh spawn on its own. Gone, with its
tests folded into `pane_free_for`'s.

`restored_screen` now drops the snapshot in both directions. Keeping the file
when it decoded to nothing left it to be re-read and re-rejected by every later
restore, and swept never, for a pane the tree still names.

Also: the module doc still said scrollback was off unless asked for, which is
what this branch reverses; and #449 landed the whole feature with no CHANGELOG
entry, so nothing told anyone that pane output now lives on disk.
2026-08-10 16:36:35 +08:00
l0ng-ai b3a66e75d0 fix(test): let the machine-tree seed keep up with a new pane field
PaneSeed grew a `shell`, but this test is unix-only, so a Windows box
never compiles it and never says so. Build the seed from `bare` and the
next field lands on its own.
2026-08-10 16:30:53 +08:00
l0ng-ai 852d3178c8 style: rustfmt 2026-08-10 16:10:58 +08:00
l0ng-ai 477d82524f feat(daemon): keep every pane's screen, without asking
`persist_scrollback` is gone, and with it the switch, its three
translations and the branches that read it. Keeping a capped tail of
each pane's output is now what the daemon does, not something it can be
asked to do.

This reverses the call made when the feature landed. The argument for
off-by-default was that the ring holds whatever the pane printed —
echoed tokens, `env` output, an agent's transcript — and that writing
that down should be the user's decision to make. What the argument
missed is when the decision gets made: the moment anyone learns they
wanted this is the moment a daemon has already died, and by then the
setting could only be turned on for next time. A feature whose entire
purpose is to survive an event nobody schedules cannot be opt-in.

The cost is real and does not go away: pane output now lives at
`<config>/scrollback/*.bin` on every machine, 0600 on unix and behind
the config directory's ACL on Windows, capped at 256 KiB per pane and
dropped as soon as no window can still ask for it.

Old configs naming the key still parse — nothing in `Config` refuses
unknown fields — so the key simply stops meaning anything.
2026-08-10 16:06:54 +08:00
l0ng-ai c138be687a fix(daemon): keep a pane's shell and its screen across a restart
Two things a pane lost when the background service stopped and started,
both of them things the tree was the only possible place to keep.

**The shell.** `PaneRecord` and `PaneSeed` carried a pane's cwd, its ssh
spec and its agent, but never what it was running. A window rebuilding a
dead pane from the tree therefore had nothing to pass and spawned on
whatever the default shell is now — so a restart turned a bash pane into
a PowerShell one, quietly and in place. The daemon resolves the override
against the config at spawn time and is the only party that knows the
answer, so it keeps it and reports it; the seed carries it too, for the
panes a window spawned itself. A handoff carries it in the blob, because
nothing on the far side of an `execve` can work out the command line of
a child it never spawned.

**The screen.** The startup sweep ran before the endpoint was listening,
which is the one moment nothing can answer the question it asks: the
registry is empty and the windows that know which screens are still
wanted cannot say so yet. A tree that failed to parse made it worse —
`read_machine` quarantines it and returns an empty `Machine`, so one bad
file took every pane's stored screen with it. The sweep now happens only
on the periodic pass, a tick later, with the registry filled in and the
tree caught up; nothing is serving a request in between. Turning the
setting *off* still clears the directory at once, because there the
promptness is the whole promise.

Two smaller ones alongside it: `restorable_pane_ids` now counts the
tree's pane list and not only the panes some tab currently stands on —
the two disagree while a window is between layouts, and being wrong
costs a file swept a tick late in one direction and somebody's terminal
in the other. And `restored_screen` drops the snapshot file *after*
deciding it was not empty, so a snapshot holding nothing is no longer
consumed by the request it could not answer.

The restore path had no end-to-end test, which is how this shipped: the
unit tests cover the file, not whether a window that reattaches is shown
anything. The new one runs a real daemon, puts a marker on a real pane,
stops the daemon, starts another, and reads the wire.
2026-08-10 16:06:54 +08:00
l0ng-aiandl0ng-ai 88bf9a5da5 feat(daemon): upgrade in place, keep pane screens across a crash, and give panes their own history (#449)
* feat(daemon): keep a pane's screen across a death nobody chose

A daemon that crashes, is `kill -9`'d, or goes down with the machine takes
every pane's replay ring with it, and the window comes back to a row of
blank shells. The processes cannot be saved that way — nothing written to
a file brings a process back — but the picture can.

The daemon now keeps a capped tail of each pane's ring under the config
directory, and a client whose `Attach` found nothing can ask, on the
`Spawn` that replaces it, for the dead pane's screen. The new pane opens
showing it, under a rule that says the shell below is new.

- periodic and dirty-only: a ring that has not moved is not rewritten, so
  an idle machine does no IO at all. Write-through would be an enormous
  amount of write amplification for a few seconds of freshness.
- capped at 256 KiB per pane, far below the ring's 8 MiB: the value of
  scrollback decays with distance from the bottom, and every byte here is
  a byte of someone's terminal on disk.
- off by default. The ring holds whatever the pane printed, including
  echoed tokens, `env` output and agent transcripts; in memory that dies
  with the daemon, and writing it down is the whole feature and the whole
  cost. Files are 0600, and turning the setting off deletes what was kept.
- dropped by relevance, not by calendar: a pane the user closed, or one no
  workspace names any more, has its file removed on the next sweep.

Restored bytes are replayed at the geometry they were written at, and are
preceded by resets — leave the alternate screen, show the cursor, restore
autowrap, clear SGR — because a snapshot is cut at the front and can begin
in the middle of any of them.

* feat(daemon): upgrade the daemon in place instead of killing every shell

Picking up a new build meant stopping the daemon, and stopping the daemon
means every pane dies: the pty master is a descriptor this process holds,
so when the process goes the slave side raises SIGHUP and takes the shell,
the agent and the half-finished command with it. That is why the update
path leaves the old daemon serving and Settings has to offer the restart
as a thing you schedule for a quiet moment.

`execve` does not have that problem. It replaces the image and keeps the
process: same pid, same children, same descriptors, same file locks. The
daemon now rewrites itself that way on `ClientMsg::Handoff` — it writes
what it knows about each pane into a blob, clears FD_CLOEXEC on the pty
masters, the blob and the singleton lock, and execs the new binary, which
picks the panes back up on the other side.

- **the seat travels on the command line, not in the blob.** The lock is
  still held by this process, so the new image must adopt the descriptor
  rather than ask for the lock again — asking would be refused by its own
  lock and it would stand down in favour of itself. A daemon that loses
  its panes is a bad afternoon; a daemon that exits leaves the machine
  with nothing serving, so that one fact has to survive an unreadable blob.
- **the blob is unlinked before it is written.** It holds every pane's
  ring, which is the output `scrollback` makes people opt into storing;
  a handoff must not be a back door for writing it to disk.
- **the exec is the last step.** Everything is staged first, so any
  failure before it costs a log line and the daemon carries on serving —
  which is what lets callers treat a failed handoff as "fall back to a
  restart" without having lost anything on the way.

Native SSH panes cannot cross — their session is cipher state in memory,
not a descriptor — so they are hung up first and the far end sees a clean
close. Windows has neither execve nor a transferable ConPTY handle, so it
keeps the stop/start path; the dialogs there still promise what they
always did, and the new copy is shown only where it is true.

Also retries flock on EINTR: a signal landing mid-call said nothing about
the lock, but was reported as "could not be evaluated", which starts a
second daemon beside the first — the split machine singleton exists to
prevent.

The end-to-end test sets a variable in the shell, hands over, and reads it
back. Nothing but the original process can answer that, and the daemon's
instance id changing while its pid does not is what says an exec really
happened.

* feat(shell): give each pane its own history when asked

Two panes running zsh with `share_history` are appending to one file and
reading each other's lines back, which is either the feature or the
problem depending on what the panes are for. Someone with a pane per task
wants Up to walk that task's commands, not an interleaving of four.

Each pane can now have its own history file instead. It is seeded from
the shell's real history, so a new pane is not blank, and what the pane
added is appended back when it closes, so nothing typed is lost — a
per-pane history that evaporated would be a way of losing commands, not
of organising them.

The seeding is done by the shell, not the daemon, and that is the only
reason it works: `HISTFILE` belongs to the user's rc file and can point
anywhere, long after the pane's environment was decided. tty7's snippet
is appended to the rc it wraps, so it runs after that decision and is the
one place the real path is known — it copies the tail, records how much it
copied, and repoints. Both shells load history after their startup files,
so the switch lands before the first line is read.

The daemon's half is a filename, a rename when a restored pane inherits
its predecessor's file, a merge on close, and a sweep for the panes a
killed daemon never got to retire.

Off by default: shared history is what a terminal has always done, and
someone who did not ask for the change would experience it as their
history mysteriously forgetting the other window. bash and zsh only —
fish and PowerShell do not keep a HISTFILE, and a shell launched with the
user's own arguments gets no snippet to repoint anything in.

* fix(daemon): store pane screens on the shutdown a restart actually uses

The periodic writer covers a death nobody prepares for and the SIGTERM
path covers a signal, but the restart the app itself performs goes through
ClientMsg::Shutdown — which killed every pty without taking a copy first.
That is the one shutdown where the panes are expected back.

* fix(daemon): leave nothing dangerous behind when a handoff fails or lands

Review findings on the in-place upgrade and per-pane history:

- A failed exec now puts back everything it had staged: FD_CLOEXEC on the
  seat and every pty master (a child inheriting the seat keeps the flock
  held past the daemon's death, so no future daemon could seat itself),
  and the SIGPIPE disposition plus this thread's signal mask, both of
  which Command::exec resets on its way to the attempt — without this,
  the still-serving daemon dies on the first client that hangs up
  mid-write.
- The adopting image restores close-on-exec on the seat and on every
  adopted master, so children it spawns later cannot hold a pty open
  past its pane, or the seat past the daemon.
- The target binary is checked before the handoff gives anything up:
  native-SSH panes are hung up on the promise that this process is about
  to be replaced, and an exec that was never going to work must not
  collect on it.
- The integration snippets raise HISTSIZE/HISTFILESIZE (bash) and
  SAVEHIST/HISTSIZE (zsh) for the pane's private history file. At their
  defaults the exit rewrite truncates the file below its own seed mark,
  which the merge-back rightly reads as "replaced under us" — silently
  losing the pane's commands for anyone with more history than the caps.
- The restart dialog's promise now binds the action: where the copy said
  "nothing is interrupted", a failed handoff is reported instead of
  silently traded for the restart that kills every pane.
- The scrollback writer checks the ring's mark before cloning it, so an
  idle pane no longer costs a full ring copy under the state lock every
  tick.

Each behavioural fix carries a test that fails without it; the history
truncation one was verified to fail with the snippet change removed.

---------

Co-authored-by: l0ng-ai <24760907+l0ng-ai@users.noreply.github.com>
2026-08-10 10:59:02 +08:00
l0ng-ai 09a653d10b fix(workspace): make the CLI and the GUI agree on what exists (#423)
Five places where a workspace, a tab or an attachment was real on one side of the socket and invisible on the other. They share a root: the GUI kept its own list of which workspaces exist (WindowViews on disk) and consulted the machine tree only for the ones already in that list, so anything created by another client was unreachable by construction.

- The switcher lists workspaces the machine holds but this client has never opened, and opening one keeps its id instead of claiming a fresh one.
- for_workspace_at hydrates whenever the machine holds tabs, so opening a workspace no longer saves an empty session over them.
- finish_hydration writes a full window back over an empty tree, which is what puts a ws rm'd workspace back under the same id.
- A deletion nothing has open is forgotten here too, instead of haunting the switcher until a restart.
- Workspace::attachment travels over the wire (minus the token that proves the hold, which stays on the connection that owns it) and is stripped in persist, so tty7 ls can name the host holding a workspace.
- tab ls / ws tree fall back through name -> agent -> cwd leaf -> process name, and tab ls grew a read-only GROUP column.
- tty7 new --open raises a window on the workspace it just made.
2026-08-09 14:18:19 +08:00
ARNOandARNO 7d86b86d35 fix(windows): prevent daemon from inheriting policy that blocks scoop junctions (#292)
Some Windows shell brokers enforce `ProcessRedirectionTrustPolicy` on what
they launch. The daemon inherited it, every ConPTY shell under the daemon
inherited it in turn, and PowerShell could then no longer traverse a
user-created junction — which is exactly what Scoop's `current` links are.
`oh-my-posh` and `fzf` died with `Shim: Could not determine if target is a
GUI app`. Windows Terminal was unaffected because its process tree never
picked the policy up.

The policy cannot be relaxed once enabled, so the fix is to not inherit it:
when tty7 detects the enforcing bit, it creates the daemon with
`STARTUPINFOEXW` and `PROC_THREAD_ATTRIBUTE_PARENT_PROCESS` naming the
interactive desktop shell, which supplies the ordinary desktop token, device
map, and mitigation policy. The Win32 code stays isolated in
`daemon/spawn/windows.rs`, and the ordinary path still runs whenever the
policy is absent — or whenever the desktop shell cannot be borrowed, in
which case tty7 logs a warning and starts degraded rather than not at all.

Because naming a logical parent makes handle inheritance follow that
process, the daemon starts with no standard handles. `daemon::server` and
the pane reader's trace line now write to stderr in a way that tolerates
that, instead of `eprintln!`, which panics on a failed write.

ConPTY exit ordering: the process-exit monitor could observe a short-lived
shell exiting before the reader had delivered its final frame, so `Exited`
reached clients ahead of the output that preceded it. The monitor now
releases the pseudoconsole and lets the reader — which reports only after
forwarding everything up to EOF — announce the death, with a bounded window
behind it for the case where EOF never arrives because a grandchild holds
the ConPTY output pipe open.

Note this changes the daemon's token on the clean-parent path: it derives
from Explorer, so an elevated tty7 starts a medium-integrity daemon.

Co-authored-by: ARNO <ArnoChenFx@users.noreply.github.com>
2026-08-02 10:59:47 +08:00
l0ng-ai 8000461706 fix(core): derive both endpoints from the config dir, publish the dir itself
Manual testing found `tty7 run`, `send`, `capture`, `procs` and `split` broken
against any normally-installed server — the CLI's entire hot path. Only the
control verbs worked.

Two endpoints, two rules. The pane socket came from the config dir; the control
socket ignored it and sat in $XDG_RUNTIME_DIR/tty7 or ~/.local/share/tty7 —
under the same basename, `daemon.sock`. So they were told apart by directory
alone, and the CLI, handed one path in TTY7_SOCKET, reconstructed the other with
with_file_name: on the default layout that returns the input unchanged. Pane
verbs dialed the control socket and the daemon hung up on them. A --config-dir
server was worse: it published the *default* control socket to the shells it
spawned, so a CLI inside an isolated instance drove a different server.

The e2e suite passed throughout because its harness set TTY7_CONTROL_SOCK
explicitly, placing both endpoints in one directory under different names — a
layout production never produces. It had removed the bug's precondition.

Now: the control socket is derived from the config dir like the pane socket
(control.sock beside daemon.sock, mirroring Windows' control.port/daemon.port,
with -control on the hashed fallback so the two cannot collide), and panes are
handed TTY7_CONFIG_DIR instead of a socket path. A CLI inherits it, so
ControlClient::connect and PaneClient::local resolve the same two sockets the
server opened, through the same functions. No second derivation to disagree.

remote_link's remote_control_socket was a third copy of the old rule, used to
locate a remote server's endpoint before connecting; it follows the config dir
too, and the env probe now reads $TTY7_CONFIG_DIR.

Drops the CLI's server-lifecycle guard: stop/start already follow the config dir
through transport::connect and --config-dir, so there is no longer a mismatch to
refuse. The e2e case that covered only `status` over a lone variable now also
runs a pane verb — the asymmetry it missed is exactly what broke.

Note: this moves the control socket for existing installs. A running pre-change
daemon will not be found at the new path, which is the honest outcome — its
control dialect is v3 against this build's v4, so reaching it only produced a
version error anyway.
2026-07-31 14:41:29 +08:00
l0ng-ai fc01b3e31f fix(client): stop timeout setup from masking the daemon's refusal
On macOS, setsockopt against a socket whose peer has already closed fails with
EINVAL. The daemon answers a bad request by writing one Error frame and hanging
up at once, so PaneSession's `set_recv_timeout(...)?` would fail before the
refusal was ever read — turning "no such pane 42", already sitting in the
buffer, into "Invalid argument".

Bounding the reply wait is an optimisation, not a correctness requirement, so it
is now best effort in both attach/observe and spawn. Nothing can hang as a
result: a closed peer returns EOF immediately, and a live peer is exactly the
case where setsockopt succeeds.

This is what made client_lib's reattach test red.
2026-07-31 13:04:51 +08:00
l0ng-ai 400d7d76b6 fix(daemon): wake displaced controllers without polling, keep busy links usable
run_stream polled pane.controls(epoch) every 200ms behind a read timeout, so
every attached pane woke its thread five times a second just to notice a
handover that may never come. The writer already learns of the handover the
instant it happens — its channel closes — so it now shuts the connection's read
side down on its way out, and the reader goes back to a plain blocking read.

SshManager::routes() reported a link as disconnected whenever its slot's
try_lock failed, which is precisely when the link is in use. The CLI's
-m <machine> refuses to route over a link it is told is down, so an actively
used connection would intermittently fail. Busy now reads as connected, matching
how SshConnection::is_alive resolves the same contention.

Also: the exit-code probe takes the child lock with try_lock, since Drop holds
it across a blocking wait(), and its window drops from 2s to 500ms — it only
needs to cover the race between pty EOF and the child becoming reapable, and
everything past that is a pane that looks frozen to every client. Observer
budget now covers status traffic and the initial replay, not just output.
Uptime is anchored where the control listener opens so a GUI-hosted server
does not report itself as freshly started.
2026-07-31 13:04:33 +08:00
thomasandClaude Fable 5 acf6ee63d8 fix(daemon): one-shot SendInput, displaced controllers close, observers get a budget
ClientMsg::SendInput (kind 55) writes to a pane's PTY without touching the
controlling subscriber, its epoch, or the size, answered by DaemonMsg::InputAck
(kind 51) or an Error for a missing or exited pane; PaneClient::send_input
wraps it. Both stay within protocol 5.

run_stream now polls its half of the socket and epoch-checks against the pane
before forwarding Input or Resize, so a controller displaced by a preempting
Attach stops writing into the shell and has its connection shut down instead of
half-open forwarding forever.

Each observer meters its queued Output through its own OutputGate; one that
lets 8 MiB pile up is pruned rather than growing daemon memory, while the
controller and the PTY never wait on it.

The pidfile reap guard accepts any legitimate daemon exe name (current exe,
tty7-app, tty7-server, tty7; .exe optional, case-insensitive on Windows), and
the agent-hooks console fast path matches tty7-server.exe and tty7.exe next to
tty7-app.exe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014JPaaZVK7rfQPKyrymzsYv
2026-07-31 10:54:53 +08:00
thomasandClaude Fable 5 5511818af9 feat(core): pane exit codes, read-only observe, procs and routed pane clients
The daemon now waits its pane child and puts the real exit code on the
Exited frame (and replays it to late subscribers of a dead pane) — the
prerequisite for tty7 run's code passthrough. The client library gains
PaneClient::observe (read-only replay+stream), PaneClient::procs, and
PaneClient::routed for reaching a remote machine's pane daemon over the
local server's ROUTE frame, mirroring ControlClient::routed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014JPaaZVK7rfQPKyrymzsYv
2026-07-31 10:18:16 +08:00
thomasandClaude Fable 5 34350018ec fix(client): PaneClient::spawn carries the workspace for TTY7_WS
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014JPaaZVK7rfQPKyrymzsYv
2026-07-31 09:50:21 +08:00
thomasandClaude Fable 5 ee69a078f0 merge: tty7_core::client — public control and pane clients
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014JPaaZVK7rfQPKyrymzsYv
2026-07-31 09:46:01 +08:00
thomasandClaude Fable 5 753d1dceee feat(core): pane observers, TTY7_* context env, and control aggregate queries
Server-side gaps for the tty7 CLI (docs/cli-design.md, the additive tier):

- Pane multi-subscriber: the single controlling subscriber keeps its
  preemption semantics, and a pane now also carries N read-only observers.
  A new pane-protocol frame `Observe { pane_id, size }` (kind 54) joins a
  connection as an observer: it gets the Snapshot replay, then
  Output/Exited/Size, and its Input/Resize is answered with an Error frame
  instead of reaching the shell. Observers never displace the GUI, never
  defer the dead-pane reap, and are pruned when their connection goes away.

- Context env injection: every locally spawned shell now carries
  TTY7_PANE (its pane id), TTY7_WS (the workspace named in Spawn), and
  TTY7_SOCKET (the control endpoint path), next to the existing TTY7
  marker. ClientMsg::Spawn gains an optional `workspace` field, carried
  by the SPAWN_OWNED frame with serde defaults.

- Aggregate queries: ControlRequest::{AgentStates, Routes, Status} with
  ReplyOk::{AgentStates, Routes, Status}. AgentStates snapshots each
  pane's live agent session state through a new Services::panes directory
  wired from the daemon's registry; Routes lists the SshManager's held
  connections with key/kind/liveness; Status reports pid, uptime from a
  process-start instant, pane count, both dialect versions, build, and
  the control socket path.

- Dialect bump: CONTROL_VERSION 3 -> 4, PROTOCOL_VERSION 4 -> 5,
  following the convention set by bed22d8 and 1792bb8/a4972d3 where every
  additive variant bumped the strict-equality handshake versions (and
  with them the tty7-server-c{control}p{protocol} install name).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014JPaaZVK7rfQPKyrymzsYv
2026-07-31 09:44:22 +08:00
thomasandClaude Fable 5 ec19525332 feat(core): extract the public client library tty7_core::client
ControlClient wraps the control-socket machinery behind connect /
connect_at / routed, adds a channel-backed events API on top of the
existing EventSink, and keeps request deadlines and blob replies.
PaneClient speaks the pane socket: one-shot list/version/kill plus
spawn/attach sessions that stream DaemonMsg and split into input and
output halves. transport grows connect_endpoint_at so a client can
reach a server by explicit endpoint file instead of the global config
dir. Integration tests drive both clients against a spawned
tty7-server --daemon in an isolated config/data dir on every platform.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014JPaaZVK7rfQPKyrymzsYv
2026-07-31 09:36:03 +08:00
l0ng-aiandl0ng-ai 8c1946d763 chore: strip every comment from the Rust sources (#268)
Removed all Rust comments -- line, block, and doc -- from the 139 tracked
.rs files with `uncomment` 3.5.1. It parses each file with tree-sitter
instead of matching text, so comment-like content inside string literals
is left alone: the JavaScript plugin source embedded in agent_hooks.rs
raw strings keeps its own `//` lines.

Left alone: Cargo.toml comments and the shell scripts under scripts/.

Co-authored-by: l0ng-ai <24760907+l0ng-ai@users.noreply.github.com>
2026-07-30 21:36:15 +08:00
a4972d32d8 feat(core): daemon-owned workspace tree — semantic ops, incremental deltas, thin clients (#260)
* refactor(daemon): share one run_daemon between tty7 and tty7-server

Extract the control-listener-plus-pane-server startup from tty7-server
into tty7_core::daemon::server::run_daemon, and point both binaries at
it. The local daemon now serves the control dialect exactly like a
remote one: one machine = one daemon, whichever binary happens to be
running it.

The bound control socket (and a bind failure) is still reported on
stderr with the historical 'tty7-server:' prefix — a headless server's
log file is off by default, and the remote_router test reads that exact
line back to prove the client derivation and the server bind agree.

* feat(core): daemon-owned machine tree with semantic operations

Add core::machine: the workspace/tab/pane tree a machine's daemon owns
outright, replacing the client-owned-schema model of the opaque record
store. Leaves hold a pane id and nothing else; every fact about a pane
(cwd from OSC 7, title, ssh spec, agent identity) lives once in the
pane registry, which is what makes revival sound: a reopened store
force-clears every live flag, so after a daemon restart the tree itself
says every leaf is awaiting revival — no client-side instance stamps or
id-reuse heuristics required.

Operations (workspace create/rename/delete/touch/set-active-tab, tab
create/close/rename/move/regroup, pane split/close/set-ratio/move/
replace) validate against the held tree, persist atomically, roll back
on a failed write, and broadcast incremental LayoutDelta events with
origin exclusion so a writer never hears its own echo. Persisted to
machine.json beside the old store's file, serde with #[serde(default)]
throughout so the daemon can keep evolving the schema, corrupt files
quarantined instead of overwritten.

* feat(control): machine-tree verbs and incremental Layout deltas

Teach the control dialect the semantic operations the machine tree
serves: MachineGet / WorkspaceTree pulls, WorkspaceCreate / Rename /
Remove / Touch / SetActiveTab, TabCreate / Close / Rename / Move /
SetGroup, and PaneSplit / Close / SetRatio / Move / Replace. Replies
carry the daemon's own tree types (a created workspace or tab comes
back whole; close operations answer the pane ids that left the tree so
the caller can kill their PTYs), and every operation broadcasts a
ControlEvent::Layout delta to every connection but the writer's — the
same origin-exclusion mechanism the record store uses, one delta at a
time instead of whole-record last-writer-wins.

The server advertises a new 'machine-tree' capability bit only when it
actually carries a MachineStore; both daemons now do, alongside the
retired opaque record store, which keeps serving unchanged while
clients migrate. Delta fan-out rides its own bounded queue and
forwarder thread per connection, so a peer that stopped reading stalls
nobody's edit; the drop-on-overflow tradeoff is documented against the
keepalive that reaps such a peer and the full pull every reconnect
starts with.

The request/reply/event enums lose their Eq derive: split ratios are
f32. End-to-end tests drive the shipped tty7-server binary over real
pipes: capability advertisement, tree ops landing in the server's own
file, dead-pane revival across a real process restart, and delta
delivery between two live clients.

* feat(daemon): pane facts flow from the pane server into the machine tree

The tree's pane records are only worth reviving from if they hold what
the machine itself observed, so the pane server now publishes into the
MachineStore the daemon serves: the reader thread reports OSC 7 / probed
cwd changes and the sniffer's agent facts (identity, native session id,
launch argv, coarse status) after each chunk that changed them, and
DeathReporter::report flips the record to live == false however the
death was noticed — that flag is the client-visible 'awaiting revival'
state, and it now comes from the process that owns the PTYs on the very
event, not only from the next restart.

The store rides a process-wide slot (installed by control_services,
same shape as the control event observer) so the three pane-spawn paths
need not thread it through; without one installed, observing is a
no-op, which keeps unit tests and tree-less servers quiet. Facts are
published outside the pane state lock and only on a real change, so the
reader's hot path pays two clones and a compare. AgentFacts.status
tightens from a free string to the existing AgentStatus enum while no
wire client depends on it.

* feat(ui): hold a supervised control link to the local daemon

The GUI now dials this machine's own daemon over the control dialect,
exactly as it does a remote one: one machine, one daemon, one control
link. The link lives in its own global rather than RemoteConnections —
inserting it there would register a wire-backed Host for this machine
(local files and git must keep going through the in-process LocalHost)
and would break the HostId::LOCAL-never-holds-a-control-connection
invariant. No routing either: the daemon's control socket is right
here, so connecting is a Unix connect plus a ControlHello.

Supervised on its own forever loop at the remote pump's cadence,
because that pump deliberately parks when the last remote workspace
closes and a purely local session is the common case. Each turn also
drains the shared control-event queue, so local pushes (Layout deltas,
Preempted) are delivered under HostId::LOCAL even with the remote pump
stopped; the observer install is shared with the remote supervisor so
whichever comes up first, reader threads never find nobody listening.
Reconnects ride the same 1/2/4/…/30s backoff a remote machine gets,
with ensure_running first — the daemon is the GUI's own child, and a
cold start legitimately races its listener.

Unix-only like the control listener it dials; on Windows the loop
compiles to a supervision no-op and the pane path is untouched.

* feat(control): attachment and takeover ride the machine tree too

WorkspaceAttach / WorkspaceDetach (and the hello-names-a-workspace
shorthand) now record their data half on whichever workspace stores the
server carries: the retired record store, the machine tree, or — on a
full daemon while clients migrate — both, since they describe the same
workspace. The behavioural contract is untouched and now survives the
record store's retirement: newcomer always wins, the displaced session
is pushed Preempted (and closed only when its link was dedicated), and
a preempted session's tidy-up detach cannot evict the usurper — the
token check lives in the tree's runtime-only attachment exactly as it
did in the store's. A server carrying neither store answers the same
refusal a store-less server always has.

WorkspaceId gains FromStr (the inverse of its Display) because the
attach verbs predate the typed tree and carry the id as a string. The
end-to-end test drives a takeover on a server serving the tree and no
record store at all, asserting the tree's own attachment record moves
with it.

* fix(core): review hardening for the machine-tree foundation

Findings from a correctness review of the new daemon-owned tree,
applied together:

- A dead pane can no longer be resurrected in the tree by its own last
  output. On Windows the exit monitor reports the death while the
  reader is still draining ConPTY's buffered bytes, and the death
  report is latched; the reader's 'output is proof of life' publish now
  asserts liveness only while the pane state still says alive.
- Delta delivery is ordered. Mutations were serialized by the state
  lock but delivered after releasing it, so one writer's deltas could
  overtake another's and leave every mirroring client on the losing
  state with no cue to re-pull. A notify-order mutex now spans each
  mutation and its own fan-out; cheap, because subscriber callbacks are
  enqueue-only by contract.
- Implicit active-tab changes broadcast. tab_create's activation and
  the close paths' heal now emit ActiveTabChanged, so a client applying
  deltas never re-implements the server's heal rule; the one
  inexpressible case (no tabs) needs no delta because it is a fact,
  not surgery.
- The coarse agent status no longer drives disk writes: it flips per
  hook event and is display-only, so it is outside the changed-facts
  gate and merely rides along when a load-bearing fact changes.
- control_services reports which stores it serves on stderr again —
  tty7-server configures no log sink, and 'no machine tree' was
  invisible exactly where it matters, on a headless box.
- The local link's first connect attempt is immediate instead of one
  backoff step late; the observation-slot test withdraws its store so
  it cannot swallow later tests' observations; and locked()'s poison
  rationale now says what is actually guaranteed.

* feat(control): let clients mint workspace and tab identities on create

A window names its workspace — in the registry, the view file, and any
operation it queues — before its first round trip completes, and the same
holds for a tab the moment the user opens it. Making the daemon the only
minter would force every client to hold its edits until a reply carried
the real id back. Ids are uuids, so a client-minted one is as unique as a
daemon-minted one; WorkspaceCreate and TabCreate now carry an optional
client id, keep it when it is free, and refuse a duplicate rather than
adopt it. Absent (older callers, tests) the daemon mints as before.

* feat(ui): windows speak semantic tree operations for every structural change

The write path of the client migration: each window now keeps a mirror of
what the daemon's tree holds for its workspace, and save_session — the
funnel every structural change already passes through — diffs the window
against that mirror and sends the recovered operations (TabCreate,
PaneSplit, PaneClose, PaneReplace, TabMove, ratio and label ops) over the
workspace's control link: the LocalLink for this machine, the machine's
RemoteConnections entry otherwise. Consecutive saves differ by exactly one
user action, so the diff recovers that action rather than re-shipping the
layout; changes no single op expresses rebuild the affected tab whole,
matching the delta contract's own granularity.

The mirror advances by running the server's own tree surgery (PaneNode's
split/remove/replace are public now), and any disagreement — a refused op,
a dropped link — resolves by one shared recovery path: drop the queue,
re-pull WorkspaceTree, re-diff. Fresh spawns are invisible until their
pane id lands; land_pane's save is when their create goes out. GUI tabs
carry a client-minted TabId, and a primed mirror re-points tabs it
recognizes by their panes, so a rebuilt window adopts the daemon's tabs
instead of churning them.

Workspace-level facts ride along: focus touches, renames, and deletions
now reach the machine's tree too, and the divider drag finally persists
the ratio it lands on (it previously reached disk only as a passenger on
the next structural change).

session.json is still written in parallel; it retires with the read-path
migration.

* feat(ui): local windows restore by asking the daemon's tree

The read path: opening a known local workspace no longer rebuilds from
session.json synchronously. The window opens empty and a background pull
(MachineGet — the workspace's structure joined with the pane registry,
which is where the revival facts live) rebuilds it the moment the daemon
answers; against the local daemon that is milliseconds, so the empty
state is effectively one frame — the same shape a remote workspace's
connect-driven rebuild has always had.

The lowering from tree to window is the revival decision: a leaf whose
pane record says live re-attaches by id, a dead one lowers to an id-less
leaf carrying the record's cwd, SSH spec and agent resume — the exact
shape that makes the existing builder spawn a successor and type the
agent's --resume. The save that follows diffs the successor against the
mirror and sends PaneReplace, spending the old record; revival needed no
op code of its own.

Restored tabs keep their daemon tab ids (SessionTab grows a never-
persisted tree_id), so the first save addresses the daemon's tabs instead
of churning them. A tree with nothing for the workspace falls back once
to the client's cached layout, whose adoption re-populates the tree
through the ordinary diff — the whole of the best-effort import.

* feat(ui): live windows apply the machine's incremental layout deltas

The pump's event drain now lands ControlEvent::Layout instead of debug-
logging it: each delta advances this client's mirror (by the same
surgery the server ran) and then the live window showing the workspace —
renames, regrouping, moves, active-tab changes and ratio drags in place;
TabCreated by building the tab and attaching its (writer-spawned, so
live) panes; TabRestructured by rebuilding the one tab while reusing the
views of panes the window already shows, because re-attaching a pane
this window holds would steal its own stream. Origin exclusion means
every delta arriving is another client's edit, and applying it to window
and mirror in one step leaves the next local diff with nothing to echo.

A delta that will not apply cleanly — a tab the mirror never heard of, a
drifted window — falls back to re-pulling the workspace and rebuilding
the window from the authoritative tree, the same single recovery path
every other failure already uses.

* feat(daemon): report panes the machine tree no longer references

With the tree now populated by clients' semantic operations, the daemon
can finally see panes nothing references. A periodic sweep reports them —
log-only, deliberately: an unreferenced pane is not proof of a leak (a
native-SSH pane opened inside a remote workspace's window runs in this
daemon while belonging to the other machine's tree), and reclaiming one
wrongly kills a session the user is looking at. The sweep's interval
doubles as a grace period: a pane is reported only after being
unreferenced across two consecutive looks, so an adoption still in
flight is never flagged. Reclamation can be layered on once the log has
shown the false-positive rate is zero.

* feat(ui): remote workspaces read and write the machine tree like local ones

Local and remote are now the same shape end to end. A remote workspace
opens empty unconditionally (connected or not) and is filled by the same
tree hydration a local window uses; the connect supervisor's landing
replaces the opaque-record refresh with it — a blinked link relinks the
pane streams and hydrates whatever opened empty meanwhile, a replaced
server process resyncs the window from the tree, whose force-cleared
live flags are what make every leaf revive. The remote picker lists
workspaces from MachineGet, deriving names from the tree the way a
local workspace derives its own; creating one lets the hydration's
WorkspaceCreate mint it on the machine; the record push, pull, refresh
(WorkspaceChanged) and remote delete paths are gone client-side.

Windows that have not yet seen their machine's tree sync additively: a
window that opened empty ahead of its pull may add tabs but never prunes
ones it has not displayed, so its ignorance can no longer read as 'close
everything' — the diff takes an explicit scope, and only hydration (or a
deliberately authoritative open, like restore-off) grants the full one.

* refactor(core): retire the client-side pane-identity defenses

The machine tree made this whole family unnecessary, so it goes rather
than lingers: daemon_instance stamps (a restarted daemon's tree says
live=false about every pane — a fact, where the stamp was a heuristic),
forget_stale_pane_ids on both layers, dedupe_pane_ids (the daemon
refuses a pane appearing twice in its tree, so there is no duplicate to
mop up client-side), the claim/record instance plumbing, and the
whole-record halves of the storage split (to_remote_json,
apply_remote_json, REMOTE_OWNED_FIELDS, CLIENT_OWNED_FIELDS, and the
store's apply_remote / remote_payload), together with their tests.

forget_pane_ids stays for now: it clears the client's cached copy, which
still serves as the one-time import fallback until the view file slims
down to pure view state.

* refactor(ui): a local daemon restart rebuilds from the tree too

The tree file survives the restart and the fresh daemon force-clears
every pane's live flag, so the resync path already expresses exactly
what the hand-rolled saved-session rebuild did: every leaf revives as a
fresh shell in its recorded cwd with its agent resumed. The pull waits
out the local link reconnecting to the fresh daemon.

* docs(core): drop a stale reference to the retired record verbs

* fix(ui): close the review findings on the tree migration

Review fixes, worst first:

- Pane ids never alias across daemon restarts: the pane registry seeds
  its counter past everything the persisted tree references. A fresh
  process minting from 1 handed new shells ids that dead leaves still
  claimed — the tree marked the wrong pane live, revival stalled forever
  on 'already part of this machine's tree', and an attach by the stale
  id stole another workspace's stream. Ids are names now, not slots.
- An empty window only licenses WorkspaceRemove once it is *informed*:
  a window whose hydration has not answered is empty because it is
  waiting, and closing or swapping it mid-pull was deleting populated
  trees. Remote workspaces also hydrate regardless of the restore
  setting — their panes are running sessions, not a saved layout, and
  the restore-off swap used to open them empty-and-authoritative and
  close every tab on the machine.
- Tabs whose panes are all still spawning are *held*, not pruned: they
  are invisible in the desired tree without being absent, and the Full
  diff was closing them (spending the records the landing spawns'
  PaneReplace needed) on every remote revival.
- A preempted window stays passive under deltas: applying the usurper's
  TabCreated/TabRestructured attached to their fresh panes and stole the
  streams they were typing into. The mirror is dropped instead; taking
  the workspace back re-pulls it whole.
- Delta TabClosed tracks the active tab by identity (closing a tab to
  the left no longer shifts focus and pushes the wrong active tab back).
- The hydrate/resync path drops the op queue like desync does, so ops
  computed against an abandoned mirror cannot drain after the snapshot.
- A rebuilt remote tab no longer matches a native-SSH leaf's *local*
  pane id against remote ids; delta-applied ratios clamp to the GUI
  band; async completions use get_mut so a forgotten window's sync state
  is not resurrected.

* feat(ui): a per-machine mirror of each daemon's tree feeds the read surfaces

The switcher, the Window menu, the title bar, the rename seeds, the
stop/delete confirmation and the liveness sweep all answered their
questions (display name, subject path, pane ids, pane count) from the
client's cached copy of the layout. The machine's tree owns the layout
now, so a new per-host MachineMirrors global holds each machine's last
pulled tree — filled by a MachineGet whenever a control link comes up
(and for free off every hydration, which already pulls the whole
machine), advanced by the same Layout delta stream the windows consume,
plus explicit notes for this client's own operations, which origin
exclusion keeps out of that stream.

The readers move over wholesale. A machine not pulled yet reads as
not-knowing rather than a stale guess: pickers show the shared fallback
for a beat (against the local daemon the pull lands within a frame),
and the pane-count prompt says the machine could not be asked instead
of counting against a cache. tree_display_name moves out of the remote
picker into the mirror as display_name_of — it was always the tree
flavour of Workspace::display_name, and now everything shares it.

This is the read-model half of retiring the client's layout cache; the
persistence shrink to pure view state follows on top of it.

* refactor(ui): client persistence shrinks to pure window views

The client file stops carrying layout. session.json's Workspace — id,
name, a whole embedded Session, geometry, open, last_active, host —
becomes WindowView { id, window, open, last_active, host } in a fresh
views.json (no migration by design; an old session.json is simply
ignored, and its panes revive from the machine tree like any daemon
restart). Everything the embedded layout used to answer already moved
to the per-machine mirror, so this deletes the write half:

- WorkspaceStore::claim answers only the id; record shrinks to
  record_geometry. claimable_session / record_session — the
  reachability-gated layout cache — go entirely, and with them the
  one-time empty-tree import in finish_hydration: with no cached copy
  there is nothing to import, and the machine answering "no tabs" is
  the layout.
- The user-set name is purely the machine's fact now. rename /
  rename_locally leave the store; the chip and switcher renames fire
  WorkspaceRename directly (tree_sync::rename_workspace), the
  WorkspaceRenamed delta needs nothing from the window because the
  mirror already applied it, and WorkspaceCreate seeds no name.
- forget_pane_ids / blank_pane_ids and the layout-derived getters
  (display_name, dominant_repo, first_cwd, pane_count, pane_ids) are
  deleted with their tests — each had grown a mirror-side twin.
- switch_workspace always hydrates: with the tree as the only layout
  source, restore-off governs what launch comes back to, not what a
  deliberate switcher pick shows.

The retired opaque record store loses its one test that asserted its
file parses as a client Workspaces document — that coupling is the
thing this migration ends, and the store itself is next to go.

* refactor(server): retire the opaque workspace record store

Clients stopped sending WorkspaceList/Get/Put/Delete when the tree
migration landed, so the coexistence scaffolding comes out:

- core::workspace_store is deleted. Attachment and the data-directory
  resolution (TTY7_DATA_DIR, XDG fallback chain) move into
  core::machine, which was already their only consumer; Attachment
  loses its vestigial serde derives (it never crosses disk or wire).
- The control dialect drops the four record verbs, the ReplyOk::Json
  payload they answered with, and the WorkspaceChanged event. Their
  serde names (and the workspace-store capability bit) are recorded as
  burned rather than reserved by any mechanism — the dialect has no
  numbered slots to hold, so a comment at each site is the guard, plus
  the handshake test asserting the bit never reappears.
- host::server loses Services.workspaces, the verb arms, the
  per-connection store subscription and its WorkspaceChanged forwarder,
  and the store half of attach/detach/teardown. Attachment data now
  lives solely in the tree: a workspace the tree does not list records
  no data half (the registry's live handles still move, so takeover
  behaviour is unchanged), and it appears the moment the workspace
  does. Services::with_workspaces/and_machine collapse into
  with_machine; control_services becomes a single match.
- The attach/takeover tests move onto MachineStore wholesale, attaching
  to workspaces created in a real tree; the record-store round-trip and
  fan-out tests go (tests/machine_tree.rs has carried the tree
  equivalents since the verbs landed), and tests/workspace_store.rs is
  deleted with the serde_json dev-dependency that existed only for it.
  machine.rs gains the two guarantees the old suite held uniquely: an
  attachment dies with its workspace structurally, and the default path
  resolution ends at the documented file.
- The GUI's dead WorkspaceChanged arm and every stale doc reference go.

* refactor(ui): rename RemoteConnections to HostLinks

Purely mechanical, plus the doc sentences that carry the model: the
table holds one control link per machine, and the local machine is a
machine like any other — its link just lives in its own global
(LocalLink) because it is in-process rather than wire-backed. The old
name framed the table as remote-only plumbing, which the tree
migration made false in spirit: local and remote windows speak the
same operations over whichever link their machine answers on.

* fix(ui): a tree-driven tab rebuild keeps the native-SSH split it cannot name

A native-SSH pane opened inside a remote workspace's window runs in
this client's own daemon and is deliberately absent from the remote
machine's tree (its local id would collide with an unrelated remote
pane). The TabRestructured rebuild therefore had no leaf for it and
dropped its view on the floor: the local session kept running,
invisible from every surface — a true orphan only the daemon's log-only
sweep would ever mention.

The rebuild now sets such leaves aside while harvesting reusable views
and appends each back as a fresh half-and-half split on the right once
the tree's own panes are built. The old split geometry is unknowable
from the delta (the tree never held it), so the appended shape is the
one a split created it in; the next save changes nothing, because the
diff already lowers a remote window without its ssh leaves.

The resync path (a delta that fails to apply, a replaced server) still
rebuilds the whole window from the tree and drops such views — that
path discards every view it has by design, and is left as a known
residual. TerminalView grows a test-only ssh-marked pane constructor so
the kept-split property is pinned by a gpui test.

* docs(core): finish pointing the last session.json references at views.json

* fix(ui): kick every local window's sync when the local link comes up

A window built while the local control link was still dialing parks as
Unprimed { dirty } — start_prime's unreachable arm leaves the retry to
"the reconnect-triggered save", but the local link supervisor never
triggered one. On a first launch (window built before the auto-spawned
daemon binds its socket) nothing else re-enters sync_window until the
next structural change, so quitting before one loses the window's
layout: the machine never heard of it.

Reproduced end-to-end on a scratch daemon: fresh launch, no user
action, quit — the relaunch came up empty. With the link supervisor
calling tree_sync::on_link_up on connect, the same launch syncs the
tree within one pump tick.

* fix(ui): read a deleted workspace's kill list before the removal blanks the mirror

delete_workspace fired WorkspaceRemove first, and fire_workspace_op folds
the removal into the machine mirror synchronously on its way out — so the
kill list stop_workspace_keeping then read off that mirror was always
empty, and 'Delete Workspace' ended zero of the sessions its confirm
prompt promised to end. The kill list is now read before the op fires,
and both destructive paths receive it explicitly so the ordering is a
signature rather than a convention.

* fix(control): bump both dialect versions and gate tree verbs on the machine-tree bit

The tree migration deleted four control verbs and added seventeen, but
CONTROL_VERSION stayed at 2 — two builds that cannot understand each
other's requests would have shaken hands as equals. It is now 3, with
the history entry the file's format asks for.

PROTOCOL_VERSION moves to 4 for the service change underneath: a
pre-tree 'tty7 --daemon' has no control listener at all, so a GUI from
this build silently adopting one connects its control link into the
void forever and every window hydrates from a tree that never answers.
The bump routes that meeting into ensure_running's existing
keep-or-restart prompt.

Clients now also consume the machine-tree capability bit before any
tree traffic: a connected peer without it (a server with no home
directory keeps serving files and panes) classifies as a distinct
'unserved' state that is logged once and skipped, instead of a refused
round trip per operation.

* fix(ui): preempted windows stay passive and take-back rebuilds from the tree

Two halves of the same takeover contract were broken.

A preempted window kept pushing: sync_window had no preemption check, so
a click on the read-only tab strip sent WorkspaceSetActiveTab against
the usurper's session, and the next save Full-diffed the stale layout —
rolling the usurper's edits back wholesale. sync_window now returns
early for a preempted workspace, and preemption itself drops the
window's queue, mirror and 'informed' licence (tree_sync::on_preempted,
shared with the delta path's existing reset).

Take Back never rebuilt: the recovery attach ran the ordinary IfEmpty
hydration, which skips any non-empty window — and a preempted window is
by definition non-empty with the pre-takeover layout. retry_now now
marks the workspace as reclaiming, and finish_attempt rebuilds marked
(or still-preempted) windows via Adopt::Replace, honouring the 'take
back re-pulls whole' promise the delta path documents.

* fix(ui): delta application survives pulls in flight

Three overlap bugs between the incremental delta stream and the full
pulls it has no ordering barrier with:

- A TabCreated straddling a pull was applied by both — the snapshot
  already carried the tab, and the delta inserted a second copy into
  the machine mirror and the window mirror, and rebuilt a second GUI
  tab whose attach stole the pane's single stream from the window
  itself. All three application sites now replace by id.

- A delta arriving while a window's prime/hydration was in flight was
  applied to the window even though the mirror side skipped it — a
  TabCreated landing in a still-empty window made finish_hydration
  read 'the user got here first' and skip adopting the tree, leaving
  the window with only the concurrently-created tab forever. Window
  application is now gated on the mirror being primed; the pull's
  snapshot carries the delta's effect.

- A prime answered after a newer cycle (hydration, desync, preemption)
  replaced it would install its stale tree over a mirror that had since
  advanced, and the next diff would re-emit the rollback as operations.
  Every cycle now stamps an epoch, and pulls landing under an old one
  are dropped.

* fix(ui): apply ratio deltas in the server's clamp band

set_gui_ratio clamped to 0.1-0.9 while the server accepts 0.05-0.95, so
another client's 0.07 arrived as 0.1 — and the next save's ratio diff
pushed the rewrite back at the machine, silently moving their divider.

* fix(core): machine-store hardening around seeds and unreadable files

- A PaneSeed entered the registry live:true unconditionally. A pane
  that died between its spawn and its adopting operation had its death
  observation dropped (note_pane_facts ignores panes the tree does not
  hold), and nothing ever flipped the record back — the leaf claimed a
  live pane forever and revival was never offered. The daemon now
  installs a liveness probe on the store (registry-backed), consulted
  at registration; without one (tests, clients) the seed is trusted.

- seed_ids_past computed max + 1, which panics a debug daemon at
  startup when the persisted tree names u64::MAX. saturating_add parks
  the counter at the ceiling instead.

- load_machine quarantined an unparseable file but not an unreadable
  one: a read failure logged, started empty, and the first mutation
  overwrote the very file that could not be read. Read failures now
  quarantine too — by rename, since a copy would need the read
  permission that just failed.

Also de-flakes the pre-existing spawn_writer test: the first write into
a freshly-closed socket can succeed before the kernel processes the
close, so the poll loop now keeps the writer fed until a write fails.

* feat(control): announce dropped layout deltas so lagged clients resync

A connection whose per-link delta queue overflowed lost an edit it will
never hear again — the server logged the drop, and the client mirrored
a tree it was no longer looking at until something else happened to
fail. The subscriber callback now flags the connection lagged, and the
layout forwarder sends the new ControlEvent::LayoutResync ahead of the
next delta it delivers (the flag is only ever set with a full queue
behind it, so the announcement never waits on a quiet tree). The client
answers by re-pulling the machine mirror and resyncing every window on
that machine — the same recovery an unappliable delta already uses,
announced instead of stumbled into. WatchOverflow is the precedent.

* fix(ui): a pure native-SSH tab is invisible to the tree, not held forever

Held means 'spawns are landing, wait before ordering' — but a remote
window's tab that is native-SSH through and through can never land: its
panes live in this client's daemon and are deliberately unnameable in
the remote machine's tree. Filing it as held made every diff return
before the ordering and active-tab passes, freezing tab order and
activation sync for the whole window for as long as the tab existed —
and a mixed tab whose last remote pane was closed kept its dead leaf on
the machine for ever, because the held id shielded the daemon tab from
the close.

Such tabs are now classified permanently invisible: not desired, not
held. Ordering resumes, and the mixed tab's daemon twin closes when its
last tree-visible pane goes. Pending leaves (a connecting spawn, an
empty slot) still read as held.

* docs(core): drop the dead instance helper, the stale title field, and two doc lies

- local_daemon_instance() lost its last caller when the client-side
  pane-identity defenses were retired; deleted.

- DaemonVersion::instance's doc pointed at Workspace::daemon_instance
  (deleted with the record store) and claimed pane ids restart from 1 —
  no longer true of a tree-carrying daemon, which seeds its ids past
  everything the tree names. Rewritten to describe what the field
  actually backs now.

- PaneRecord::title claimed to label panes awaiting revival, but no
  code ever wrote it: the pane's title is a live foreground-process
  query at PaneInfo time, not state the facts path observes. The field
  is deleted (serde-compatible: unknown fields are ignored on read) and
  the decision recorded where it lived; revival labels derive from cwd
  and agent.

* fix(ui): converge the tree after adopting a delta-created tab

Adopting a TabCreated delta whose pane is dead on arrival attaches
nothing and spawns a fresh pane under a new id — and nothing on the
delta path saved afterwards, so the tree kept the dead leaf: other
clients saw a dead tab, and a relaunch would spawn a second successor
beside the leaked first. Reproduced end-to-end (external client creates
a tab with an unspawned pane; the GUI adopted it and the tree never
learned the successor's id).

One sync_window after a clean apply closes it: free when window and
mirror agree (the diff is empty), and exactly the PaneReplace that
spends the dead record when adoption had to spawn.

* fix(core): review follow-ups on the daemon-owned tree

Nine findings from a review pass over the branch. One commit because
they cross the same files, and splitting them would leave an
intermediate that does not build on Windows.

- A dropped delta announced a LayoutResync and then delivered the
  backlog behind it. The queue is FIFO, so everything still in it is
  *older* than the gap: the peer re-pulled on the notice and was then
  walked back through history it had already left — TabRestructured
  restoring the shape a tab used to have, with window and mirror
  agreeing on the stale answer so nothing recovered a second time. The
  forwarder now drops the superseded queue and sends the resync in its
  place.

- Pane facts persisted the whole document, with an fsync, from the PTY
  reader thread — once per OSC 7, so once per prompt per pane — while
  holding the lock that orders every other client's edits. A shell
  looping over directories was a write per iteration. Observations
  (pane facts, workspace_touch) now take Persist::Soon: the delta still
  goes out at once, the file catches up within FACT_FLUSH_INTERVAL, and
  the daemon flushes on the way out. The layout itself is never
  deferred.

- An ordinary output chunk paid two AgentFacts clones and a
  clone-to-compare for facts it could not have changed. Gated on the
  signals that can move one, and the compare no longer clones.

- machine.json was created 0644, naming every workspace's directories,
  the SSH user and host of every native-SSH pane, and each agent's
  session id. It is written owner-only from the first instant the final
  name exists, and a second corruption no longer overwrites the rescue
  copy of the first.

- Windows had no control listener, so on the one platform where the
  tree is the only layout store, tabs did not come back at all. It now
  serves the dialect over the transport its pane socket already uses: a
  loopback listener whose port and 256-bit token live in a user-private
  control.port beside daemon.port — its own token, not the pane
  endpoint's — refusing to rebind over a live one, since binding is
  what writes the marker. run_daemon and the GUI's local link are one
  code path again.

- Workspace names and paths came only from the machine's mirror, so a
  laptop shut since Friday listed every row as "Untitled" with a blank
  subtitle, in the picker whose whole job is offering workspaces on
  machines that are asleep. WindowView carries the label and subject
  the machine last gave, stamped on save and on detach; the tree still
  wins whenever it answers.

- liveness_of read "the mirror has not been pulled yet" as Stopped,
  which tells the user their sessions are gone on the strength of our
  own ignorance. Unknown is what that state is for.

- A WorkspaceRemove that never reached its machine was a debug line,
  though the client had already forgotten the workspace. It is now a
  warning that says what was left where.

- MachineMirrors::install landed a pull without a repaint; the two tests
  the record store's retirement took with it (a closed connection stops
  being a subscriber, concurrent connections can all write) are back
  against the tree; and CHANGELOG records the migration's one-time
  layout loss and the Windows gap this closes.

Suites green: tty7-core 675, tty7 819, tty7-server 9/5/3/3/51, fmt and
clippy clean. The Windows listener is unverified by a compiler here — a
C dependency in the tree blocks cross-checking from macOS — so CI's
Windows job is its first build.

---------

Co-authored-by: l0ng-ai <24760907+l0ng-ai@users.noreply.github.com>
Co-authored-by: thomas <thomas@thomass-Mini.lan>
2026-07-30 12:25:38 +08:00
l0ng-aiandl0ng-ai bed22d899e Keep workspaces whole: remote reopen/restart recovery, and cross-workspace restore guards (#257)
* feat(remote): keep a remote workspace whole across reopens and restarts

Reopening a remote workspace — or coming back to one whose `tty7-server`
had been replaced — landed on a screen of `tty7 — disconnected` panes with
their coding-agent conversations gone. Several independent holes added up
to that; this closes them together, and picks up the surrounding work the
same session produced.

**Telling a restarted server from a blinked link.** `ControlHelloOk` now
carries an `instance` minted once per server *process*. Nothing else in
the handshake changes across a restart — `build` and both dialect numbers
survive it — so a reconnect had no way to know its `pane_id`s were dead.
It does now: a different instance rebuilds the window from its layout
(same tabs and splits, fresh shells in the saved cwds) instead of
re-attaching to a process that is gone. An absent instance means *unknown*
and is never read as a restart.

**An attach can now fail.** `Attach` has no synchronous reply, so the
client returned `Ok` unconditionally and the daemon's `Error` frame was
read much later by the reader thread, which has no arm for it — the pane
then landed in the *link is down* state instead of falling back to a fresh
shell. The client now reads far enough into the reply to classify it on
the kind byte (the snapshot behind it can be megabytes) and hands those
bytes to the reader thread, so a successful attach loses none of its
replay. Local and remote attaches get different waits: the local one is on
the UI thread.

**The agent session survives to be resumed.** `TerminalView` raises
`AgentSessionChanged` when the pane's agent reports a new native session
id, so the layout on file catches up instead of waiting for the user to
happen to open a tab. A pane that is still connecting now carries its
agent through `PendingSpawn` — a save landing in that window used to write
`agent: null` over the record — and `land_pane` sends `--resume` when the
attach turned out to need a fresh shell.

**Ending sessions says so on file.** "End Sessions" kills the panes and
then drops their ids from the record, pushing the cleared layout to the
machine that owns it (design §10: the remote's copy wins, so a local-only
clear would be undone by the next open — the open this exists for).

**The new-tab dropdown lists the window's machine.** `Host::shells` and a
`Shells` control request (dialect v2) make the "+" menu a property of the
machine the window is bound to. A remote window filled from this
computer's `/etc/shells` offered `/bin/zsh` on a box whose zsh is
elsewhere, and every pick failed to spawn.

**An install reports its bytes.** The download and the SFTP upload each
report progress, relayed to the client over the routed connection as a
`RoutePrompt::InstallProgress`, and painted as a bar under the machine's
row in the switcher. ~8 MB across two hops behind the word "connecting…"
was indistinguishable from a hang.

**The installer compares dialects, not version strings.** `tty7-server
--protocol` prints what a binary speaks without starting it, so a connect
adopts an already-running server it can talk to rather than prompting
about a build difference and uploading 8 MB the machine did not need.

**Switcher.** A machine's `⋯` menu holds "New Workspace" (it was a row
under every machine, pushing the list a quarter of a card down) and a new
"Disconnect", which drops the connection and leaves the windows open and
read-only. The suspension lasts exactly as long as that machine has a
window on it.

Also drops three design/contract docs for the now-shipped remote-workspace
work.

* fix(session): stop one workspace's panes from being restored into another

A restart put a copy of one workspace's seven tabs — cwds, layout and
recorded agent sessions — in front of another workspace's own tabs, and
auto-resumed every one of those agents a second time: six `claude
--resume <id>` pairs running in parallel against the same conversations,
one set per window. The record-level corruption that seeded it is still
unattributed, but every mechanism that let it propagate, amplify, or go
unnoticed is closable, and this closes them.

**Panes now know their owner.** `Spawn` can carry the workspace the pane
is created for; the daemon stores it immutably and reports it in
`List`'s `PaneInfo.owner`. Restore refuses to re-attach a pane another
workspace owns (`pane_attachable`) — before this, a saved id landing on
somebody else's live pane attached silently, which is how one window
could pick up another's shells. The field rides a new `SPAWN_OWNED`
frame with a struct payload (the legacy spawn payloads are positional
tuples an old daemon cannot grow), gated on a new `pane-owner` feature
string: a client only sends it to a daemon that advertises it, so the
legacy kinds stay byte-for-byte what old daemons expect. A pane with no
recorded owner stays attachable by anyone — that is the pre-field
behavior, not a new risk.

**Saved pane ids are bound to the daemon process that issued them.**
`DaemonVersion` now carries an `instance` minted once per process (the
local twin of the control hello's), the GUI caches it at the
`ensure_running` handshake, and each local workspace records it as
`daemon_instance` beside its layout. Claiming a workspace whose ids came
from a different instance blanks them first: daemon pane ids restart
from 1, so after a reboot every saved id points at whatever unrelated
shell holds the number now, and the aliveness check cannot tell a
survivor from a squatter. A blank on either side means "cannot tell" and
never trips it. Unlike the duplicate-claim case below, this path keeps
the agent resume — the pane is genuinely gone with its daemon, and the
fresh shell resuming the conversation is the feature.

**A duplicate claim loses its agent resume along with its pane id.**
`dedupe_pane_ids` kept the loser's layout *and* its
`agent_session_id`, so the blanked leaves took restore's spawn-fresh
path and auto-typed `claude --resume` for conversations the winning
workspace's panes were still running — the doubling above. The winner
keeps the panes and the resume; the loser keeps only cwds.

**Cross-workspace saves are caught at the write.** Every terminal view
remembers the workspace whose window created it, and `save_session`
logs an error naming both ids if a window ever records a pane created
for a different workspace — the tripwire for the still-unattributed
seed corruption, so a recurrence is caught in the act instead of
reconstructed from `session.json` archaeology days later.

Wire compatibility both ways: `PaneInfo.owner`, `DaemonVersion.instance`
and `Workspace.daemon_instance` are `#[serde(default)]` struct fields
(old peers' JSON decodes, new fields are ignored by old readers), and
`SPAWN_OWNED` is feature-gated as above. `daemon_instance` is
client-owned in the design-§10 storage split — it names the local
daemon, and the field-census test pins the classification.

* fix(session): resume the agent when a local pane dies mid-restore

`session_to_pane` decided whether to send a coding agent's `--resume`
from `restore.is_none()` — i.e. from whether the pane looked alive when
the restore started. But `alive_panes_on` runs one `List` at the top of
the restore, while the attaches happen per leaf afterwards. A pane that
exited in between failed its attach, fell back to a fresh shell inside
`spawn_shell_terminal_in`, and then landed in the `restore.is_some()`
arm: an empty shell with its conversation dropped.

`ShellParts.restored` already answers this exactly, and the remote path
already reads it in `land_pane`. Carry it onto `TerminalView` so the
synchronous local path can read it too, and branch on that instead of
re-deriving the answer from a set that may be stale by the time it is
used.

No behaviour change on the paths that were already correct: a view that
was never restoring anything reports `restored: false`, which is the
same answer `restore.is_none()` gave them.

* fix(remote): check the server instance against the record, not just memory

A remote workspace's pane ids were only guarded against server restarts
by `RemoteLinks::instances`, an in-memory map. On the first connect after
the client starts, every machine is a first sighting, so `server_restarted`
answers false — and a `tty7-server` that was replaced while the client was
closed sails straight through. Its pane ids restart from 1, so the saved
ones now name unrelated shells, and the reconnect attaches to them: the
exact id-reuse failure the local side already guards against.

`Workspace::daemon_instance` was local-only for the stated reason that a
remote server's identity is tracked live per connection. That tracking is
correct but not sufficient — it cannot survive the client restart that
makes the question worth asking.

So the field now means the same thing on both sides: which process minted
the pane ids in this record. `WorkspaceStore::serving_instance` picks the
local daemon or the far machine's server depending on the workspace, and
`finish_attempt` compares it per workspace before deciding to re-attach or
rebuild. It stays client-owned: it records what *this* client last saw, so
two clients on one remote workspace each keep their own and neither may
overwrite the other's.

An unreachable machine still records nothing, which is what keeps a good
stamp from being erased with `None` — that would disarm the next check.

Also in these three files: the §N references to the deleted design docs,
cleaned up as part of the sweep in the following commit.

* docs: drop the references to the deleted design documents

The three documents this branch removed were cited ~280 times: `design
§10`, `contract §8`, `§17` and friends in comments, five references by
file path in code and manifests, five in CI workflows and one in the
release skill. Every one of them now points at nothing.

Rewritten rather than merely stripped, because most were not decoration:
"design §10 makes the remote's `workspaces.json` the authority" becomes a
statement in its own right, and the several that carried a Chinese phrase
from the document as their justification say the same thing in English
instead. Where the reference was purely parenthetical it is simply gone.

Not touched: `PRD §7.1`, `brief §8` and the like, which name documents
this branch did not remove and were already external before it, and the
`RFC 4648 §10` test-vector citation, which is a real specification.

The `host boundary` CI job loses `(§10.6)` from its name. It is not one of
the required checks, so branch protection is unaffected.

---------

Co-authored-by: l0ng-ai <24760907+l0ng-ai@users.noreply.github.com>
2026-07-29 19:15:19 +08:00
l0ng-ai 0ec8e050ad fix(workspaces): make a handover atomic and stop two stores clobbering one file
The takeover moves two things — the `WorkspaceStore`'s record and the
server's `AttachRegistry` handle — and each was internally locked, which
is not the same as the pair moving together. Two clients attaching one
workspace at the same instant could each win a different table, after
which the store named a session the registry had already evicted and no
`detach` could clear it: the workspace reported a takeover against a
client that had disconnected hours ago. Both moves now happen under one
handover lock, dropped before the displaced client is written to so a
peer that has stopped reading still cannot hold up the next attach.

`WorkspaceDelete` had the same split with no race needed at all: the
store dropped its attachment and the registry kept its handle, so the
next client to attach that id evicted a session nobody displaced — and,
that entry being dedicated, closed its whole link.

Two `tty7-server --stdio` sessions arriving while no daemon was up each
served in-process, each with its own store over the one file. `persist`
writes the whole document, so the second to save silently dropped the
first's changes, and their separate registries made takeover a no-op
between them. The probe path now starts the daemon and bridges to it —
the rule `bridge_panes` already follows one dialect over — and the store
re-reads when the file has moved underneath it, which covers the cases
where two writers are deliberate.

`MAX_RECORD_BYTES` and `MAX_WORKSPACES` did not bound their product:
seventeen maximal records put the array past `MAX_FRAME`, after which
every `WorkspaceList` was unencodable and every client showed an empty
list. The document is now bounded at the save, and only when growing, so
an over-large file can still be deleted back under the limit.
2026-07-28 21:54:23 +08:00
l0ng-ai 26f3a73f58 fix(ci): gate the tty7-server test suites that need --stdio on Unix
`--stdio` is refused on Windows by design, and the control socket it
probes for is Unix-domain, so `stdio_conformance` and `workspace_store`
join `remote_router`/`routed_pane` in carrying a file-level `cfg(unix)`.

`cli.rs` keeps its argument-handling cases everywhere — `--version`,
`--help`, `agent-hook` and the usage error say nothing about transports
— and gates only the bridge and probe cases, which spawn a `--stdio`
child or stand up a listener.
2026-07-28 19:39:19 +08:00
l0ng-ai 208454e202 feat(remote): remote workspaces — a window that is one machine
Split the framework-free half of tty7 into `tty7-core` and add a headless
`tty7-server` built on it, so a workspace's filesystem, git and session state
can live on another machine while the GUI stays where it is.

- `crates/tty7-core`: wire protocol, session daemon, PTY, native SSH engine and
  the domain model, with no gpui dependency. Module paths are unchanged.
- `crates/tty7-server`: the same daemon with no GUI attached, linked fully
  static against musl and pushed onto the remote box. One dependency, on
  purpose — a second one the GUI also needs belongs in core.
- `Host` trait + `HostId`/`HostRegistry`: every fs/git/watch call a workspace
  makes goes through the machine it belongs to. `LocalHost` answers on this
  box, `RemoteHost` over a routed control connection.
- `ui::host_ops`: the GUI's single door to a `Host`. Host calls block, so all
  of them run on the background executor with the result landed on the UI
  thread; de-duplication, staleness and error reporting live here rather than
  at each call site. Enforced by a CI grep.
- Connect flow: home page → pick a configured SSH host → the machine's own
  workspace list → a window bound to one workspace on it. Workspace switcher
  groups by machine, this computer included.
- CI: static musl builds of `tty7-server` for x86_64/aarch64 via
  cargo-zigbuild, a host-boundary grep, and version stamping factored out of
  the nightly workflow. Both new jobs are non-required so branch protection
  does not wedge open PRs.

Design and the interface contract it was built to are in
`docs/2026-07-27-remote-workspace-{design,impl-contract}.md`.
2026-07-28 10:59:46 +08:00