mirror of
https://github.com/l0ng-ai/tty7.git
synced 2026-10-04 08:01:59 +00:00
A daemon can survive quit-and-stop with its endpoint unlinked and its pidfile gone while still holding the singleton seat (#667). Every later launch then spawns a daemon that stands down against the lock and times out red, and nothing on the machine can recover: stop answers "not running", ensure_running reaps only through the pidfile, and flock cannot say who the holder is. Two roads led there, and both are closed: - The reap identified a daemon by proc_pidpath alone, which fails outright for a live process whose binary was deleted — every nightly update replacing the installation. The identity check now falls back to the kernel's comm name (proc_name on macOS, /proc/pid/comm on Linux, both recorded at exec and immune to deletion), strips Linux's " (deleted)" marker, and — decisively — no longer deletes the pidfile of a live process it cannot identify: the record was the only handle left on the survivor. - When the pidfile is gone entirely, the pid the claimant now writes into daemon.lock at claim time is the handle of last resort. The lock file is never deleted and holding the flock is the definition of being the server, so while the seat is held its content names the holder; stop() and the reap fall back to it, and a confirmed reap clears the record (only under a momentarily-free seat) so a stale number cannot outlive its process. Unix-only: the Windows seat is share_mode(0), unreadable while held. Every road back now clears a stranded seat, not just the GUI's: ensure_running's stale cleanup is factored into spawn::reap_stranded, tty7 server start runs it too, and tty7 server stop no longer takes "nobody answered" for "nothing to stop" when the seat is still held. A short grace keeps the reap away from a daemon that is merely mid-handoff or mid-startup — where health is an answered handshake, never a bare connect: a wedged daemon's listener still completes connections out of the kernel's backlog. The startup-timeout errors name the seat-holding pid, with the kill advice identity-gated so a stale record never tells anyone to kill an innocent process. Two liveness corrections round it out: a zombie now reads as dead — it answers kill(pid, 0) like the living but holds no lock and no image, and no signal can end it, so counting it alive spent both reap timeouts on a corpse (the GUI never waits on the daemons it spawns, so crashed daemons are zombies as a rule) — and stop() only pays the process-exit wait for a shutdown it actually delivered, instead of watching an unreached survivor not move for five seconds. The guard tests were each verified to fail against the behavior they pin (fallbacks, the handshake criterion, the grace, and the wait gate removed by mutation) before being trusted green; the zombie probe semantics (proc_pidinfo failing for a zombie that still answers signal 0) were measured, not assumed.
159 lines
5.7 KiB
Rust
159 lines
5.7 KiB
Rust
//! Guards for #653: a daemon that lingers after unlinking its endpoint must
|
|
//! stay findable and reapable, or its singleton lock makes every later launch
|
|
//! stand down and time out red with no way back short of `pkill`.
|
|
//!
|
|
//! Unix-only: both tests drive the daemon through unix sockets and signals.
|
|
|
|
#![cfg(unix)]
|
|
|
|
use std::path::Path;
|
|
use std::process::{Child, Command, Stdio};
|
|
use std::time::{Duration, Instant};
|
|
|
|
use tty7_core::client::PaneClient;
|
|
use tty7_core::daemon::protocol::ClientMsg;
|
|
|
|
const READY_WITHIN: Duration = Duration::from_secs(30);
|
|
const EXIT_WITHIN: Duration = Duration::from_secs(10);
|
|
|
|
fn spawn_daemon(dir: &Path) -> Child {
|
|
Command::new(env!("CARGO_BIN_EXE_tty7-server"))
|
|
.arg("--daemon")
|
|
.arg("--config-dir")
|
|
.arg(dir)
|
|
.env("TTY7_DATA_DIR", dir)
|
|
.env("TTY7_CONTROL_SOCK", dir.join("control.sock"))
|
|
.stdin(Stdio::null())
|
|
.stdout(Stdio::null())
|
|
.stderr(Stdio::null())
|
|
.spawn()
|
|
.expect("start tty7-server --daemon")
|
|
}
|
|
|
|
fn await_ready(dir: &Path) {
|
|
let endpoint = dir.join("daemon.sock");
|
|
let deadline = Instant::now() + READY_WITHIN;
|
|
while PaneClient::at(&endpoint).version().is_err() || !dir.join("daemon.pid").exists() {
|
|
assert!(
|
|
Instant::now() < deadline,
|
|
"tty7-server did not open its endpoint within {READY_WITHIN:?}"
|
|
);
|
|
std::thread::sleep(Duration::from_millis(50));
|
|
}
|
|
}
|
|
|
|
fn send_shutdown(endpoint: &Path) {
|
|
use std::io::Write as _;
|
|
let mut stream =
|
|
std::os::unix::net::UnixStream::connect(endpoint).expect("connect to the daemon");
|
|
ClientMsg::Shutdown
|
|
.encode(&mut stream)
|
|
.expect("send Shutdown");
|
|
stream.flush().expect("flush Shutdown");
|
|
}
|
|
|
|
fn await_exit(child: &mut Child) -> std::process::ExitStatus {
|
|
let deadline = Instant::now() + EXIT_WITHIN;
|
|
loop {
|
|
if let Some(status) = child.try_wait().expect("query the daemon's state") {
|
|
return status;
|
|
}
|
|
if Instant::now() >= deadline {
|
|
let _ = child.kill();
|
|
let _ = child.wait();
|
|
panic!("the daemon did not exit within {EXIT_WITHIN:?} of Shutdown");
|
|
}
|
|
std::thread::sleep(Duration::from_millis(50));
|
|
}
|
|
}
|
|
|
|
/// The pidfile must outlive the daemon's own cleanup: once `daemon.sock` is
|
|
/// unlinked it is the only name anything has for the process, and an exit
|
|
/// that stalls after the cleanup (#653) is findable through it — deleting it
|
|
/// there is what turned a lingering process into a permanent lockout.
|
|
#[test]
|
|
fn a_clean_shutdown_keeps_the_pidfile_until_the_process_is_gone() {
|
|
let dir = tempfile::TempDir::new().unwrap();
|
|
let mut child = spawn_daemon(dir.path());
|
|
await_ready(dir.path());
|
|
let pid = child.id().to_string();
|
|
assert_eq!(
|
|
std::fs::read_to_string(dir.path().join("daemon.pid"))
|
|
.unwrap()
|
|
.trim(),
|
|
pid,
|
|
"the pidfile names the running daemon"
|
|
);
|
|
|
|
send_shutdown(&dir.path().join("daemon.sock"));
|
|
let status = await_exit(&mut child);
|
|
|
|
assert!(status.success(), "clean shutdown exits cleanly: {status:?}");
|
|
assert!(
|
|
!dir.path().join("daemon.sock").exists(),
|
|
"shutdown unlinks the endpoint"
|
|
);
|
|
assert_eq!(
|
|
std::fs::read_to_string(dir.path().join("daemon.pid"))
|
|
.unwrap()
|
|
.trim(),
|
|
pid,
|
|
"the pidfile survives the daemon's cleanup; the reap in spawn::stop \
|
|
and ensure_running deletes it once the process is confirmed gone"
|
|
);
|
|
}
|
|
|
|
/// `spawn::stop` must reap a daemon that stopped listening but never exited
|
|
/// (#653) — and without the shutdown wait: only a shutdown that was actually
|
|
/// delivered earns `PROCESS_EXIT_TIMEOUT`, the time an *asked* daemon gets
|
|
/// to finish exiting. To a survivor stop() could not even connect to, that
|
|
/// wait was five seconds of watching nothing move (#667); the reap's own
|
|
/// SIGTERM window is all the grace such a process gets.
|
|
///
|
|
/// (The companion ordering — the pidfile already gone when the reap needs a
|
|
/// pid — is pinned by the `daemon_reap` suite through the seat record.)
|
|
#[test]
|
|
fn stop_reaps_an_unreachable_daemon_without_the_shutdown_wait() {
|
|
let dir = tempfile::TempDir::new().unwrap();
|
|
tty7_core::core::config::set_config_dir(dir.path().to_path_buf());
|
|
let child = spawn_daemon(dir.path());
|
|
let pid = child.id();
|
|
await_ready(dir.path());
|
|
|
|
// Collect the child the moment it dies: outside tests the daemon is
|
|
// nobody's child, and a zombie would read as alive to the reap's
|
|
// liveness poll.
|
|
let waiter = std::thread::spawn(move || {
|
|
let mut child = child;
|
|
child.wait()
|
|
});
|
|
|
|
// The #653 state, as stop() meets it: the endpoint is gone before stop()
|
|
// can ask for a shutdown, and the process lives on.
|
|
std::fs::remove_file(dir.path().join("daemon.sock")).unwrap();
|
|
|
|
let stop_started = Instant::now();
|
|
tty7_core::daemon::spawn::stop();
|
|
|
|
// Well under spawn.rs's PROCESS_EXIT_TIMEOUT (5s): a stop() that pays
|
|
// that wait for a shutdown it never delivered has regressed.
|
|
assert!(
|
|
stop_started.elapsed() < Duration::from_secs(4),
|
|
"stop() spent {:?} on a daemon it never reached — the exit wait must \
|
|
be earned by a delivered shutdown",
|
|
stop_started.elapsed()
|
|
);
|
|
let deadline = Instant::now() + Duration::from_secs(2);
|
|
while unsafe { libc::kill(pid as libc::pid_t, 0) } == 0 {
|
|
if Instant::now() >= deadline {
|
|
unsafe { libc::kill(pid as libc::pid_t, libc::SIGKILL) };
|
|
panic!("stop() left the daemon (pid {pid}) holding the singleton lock");
|
|
}
|
|
std::thread::sleep(Duration::from_millis(50));
|
|
}
|
|
waiter
|
|
.join()
|
|
.unwrap()
|
|
.expect("collect the reaped daemon's exit");
|
|
}
|