* fix(codex): stop a transient filesystem error from logging out the active account
A single unreadable read of a managed Codex home's ownership marker cleared the
user's active account selection, permanently. On Windows any exclusive lock —
Defender real-time scanning, a backup agent, a sync client — makes every read of
that marker fail with EBUSY, and the background rate-limit poll runs every 15
minutes plus once at every app start.
Root cause: the ownership gate answered two very different questions through one
channel. "This home is not ours" (a successful observation that failed a trust
check) and "we could not read it" both surfaced as a throw, which the caller
flattened to null, which three call sites took as proof the home was
untrustworthy and wrote activeCodexManagedAccountId: null.
Refusing to USE an unverified home is correct. Erasing the user's account
selection because a file was briefly locked is not.
The gate now returns a tri-state verdict. `untrusted` comes only from a proven
trust failure or a definitive ENOENT/ENOTDIR where absence is itself the
verdict; every other filesystem exception is `indeterminate`. Only `untrusted`
may touch persisted state.
Because `null` already meant "fall through to the system default" on both the
launch and poll paths, not-clearing on its own would have run a DIFFERENT
account behind a UI still showing the selected one. So the refusal needed real
channels rather than a sentinel:
- the poll returns an explicit skip; returning null would not have skipped at
all, since the fetcher maps null to ~/.codex and would have spawned a
token-refreshing app-server inside the user's real credential home
- pane launch throws a typed temporary-unavailability error that both PTY
implementations convert into a clean refusal with a retry message, including
the re-resolution after the async auth-readiness wait
- automatic session resume resolves the selected home eagerly, so an unreadable
account can no longer be silently replaced by another one in the ranking
- config-sync status reports a distinct managed-home-unavailable stall instead
of "synced", with a bounded renderer retry so it clears on its own
Also fixes the ticket's second symptom. The status bar's Sign in button called a
re-auth that captured the selection before login and restored it after, so
re-authenticating a deselected account restored `null` — a successful login that
left the account inactive, with no success toast to distinguish it from failure.
It now activates the account it just signed in, but only when the pre-login
selection was empty, so it cannot silently switch accounts for multi-account
users, and it runs the same restart prompt an explicit switch does.
No retry or grace window inside the synchronous gate: it runs on the Electron
main process in a loop over accounts, so a sleep there would freeze the UI.
Recovery is simply the next readable evaluation.
The WSL lane has the same class of defect, including one path that deletes a
credential mirror. It is pre-existing, unreachable from these host code paths,
and deliberately left for its own change; the host clearing sites cannot reach a
WSL account because getSelfContainedManagedHostAccount excludes them.
Fixes STA-4422
* test(codex): cover pending reset home ownership