mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 16:02:37 +00:00
2d36a09cddebb23e3f32caa220e0bdcfb9be5bc1
10730
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2d36a09cdd | test(ai-vault-search): compare the whole query, so a quoted operator value survives | ||
|
|
b66b0ad431 | fix(ai-vault-search): highlight only the marks FTS5 inserted, not the text's own | ||
|
|
0e3bc9281f | fix(ai-vault-search): read a fractional or negative cursor generation as malformed | ||
|
|
4eb59e32fa | fix(ai-vault-search): keep a scope nothing could key from widening the search | ||
|
|
1dc747ce83 | test(ai-vault-search): make the trigger-restore test exercise the trigger it drops | ||
|
|
a974a39592 | docs(ai-vault-search): say which delete paths fire the orphan-reclaim trigger after a replace cuts loose | ||
|
|
4d485b28dd |
refactor(ai-vault-search): keep the scope's expression with the other expressions
Making the typo repair ask its questions in the search's own scope put an import from retrieval into it, and retrieval already owns the repair — a cycle the native audit catches. `scopedExpression` belongs beside `phraseExpression`, `andExpression` and `orExpression` anyway: it builds a MATCH expression, and two callers now need it. The bm25 weights stay in retrieval, where the SQL that uses them is. |
||
|
|
cf5860373a |
test(ai-vault-search): make each snippet mark mechanism answer for itself
Two mechanisms landed together and hid each other: choosing a column by comparing a marked rendering against an unmarked one, and marking with private-use code points instead of `[[`. Either alone fixed the bracket repro, so neither had a mutation against it — the same masking the snippet's two column guards had a round ago. They do different jobs, so both stay and each gets the test that needs it. A transcript holding a private-use code point of its own is what the comparison is for; agent output carries Nerd Font glyphs from that block. A snippet past the character ceiling with a bracket after its last real mark is what the private-use marks are for, because the truncation has to find that mark by searching the text. The two one-line fixes get honest framing rather than a mutation neither can have. A lone surrogate is not a token character, so the planner drops it either way and the safe cut is hygiene. And no SQLite this stack runs refuses 1,100 bound ids — 32,766 has been the floor since 3.32 — so the batch is about owning the ceiling here rather than rescuing a reachable failure. |
||
|
|
a5d178b9af |
docs(ai-vault-search): price the repair rung, and record what is left open
Typo repair is the one rung whose cost tracks the vocabulary rather than the result, and it only runs for a term the scope has no posting for. Measured over 1.6 M distinct terms: 10 ms for one unknown term, 387 ms for a 480-character query of thirty-nine of them. The scoped-count fix made that cheaper rather than dearer, from 737 ms, because ordering the vocabulary scan by term drops the sort `doc DESC` needed and the counts it adds are at most eight bounded probes per prefix. A cap on unknown terms per query is a follow-up in the split plan, with the five other items the final review raised and did not route. |
||
|
|
64aff99396 |
fix(ai-vault-search): cut a query on a code point and bind ids in batches
Two small ones from the review's not-routed list. `query.slice(0, 512)` can land between the halves of a surrogate pair, leaving a lone half that matches nothing and that a caller cannot echo back. The reader already has `sliceAtCodeUnitLimit` for exactly this. `loadSessions` bound one parameter per candidate id in a single statement. The list is as long as the candidate limit, the tuning doc invites a host to raise that limit, and SQLite's `SQLITE_MAX_VARIABLE_NUMBER` is 999 on builds older than 3.32 — so one settings change away from `too many SQL variables`. Read in batches of 500, leaving room for the filter's own bound values. |
||
|
|
7c1f0874a6 |
fix(ai-vault-search): tell a highlight from a transcript that contains brackets
The snippet builder asked each of a row's four columns for a marked snippet and showed the first whose text contained `[[`. Transcripts contain `[[`: a bash `if [[ -f … ]]`, numpy's `[[1, 2], [3, 4]]`. A row matching only in tool output was shown its user turn instead, with nothing highlighted in it, and the any-column fallback an identifier-only match depends on was unreachable behind the same collision. Whether a column matched is now the difference between two renderings of the same text: `snippet(…, MARK, MARK, …)` beside `snippet(…, '', '', …)`. Content cannot forge a difference between those two, because it is the same content either way. The marks FTS5 inserts are private-use code points, rewritten to the public `[[` and `]]` once, at the end. That is for the other decision that has to tell a mark from content: the truncation refuses to cut between an open mark and its close, and a transcript's own bracket used to move that cut. |
||
|
|
739197e55f |
fix(ai-vault-search): fence the rows a purge reclaims after it cuts a session loose
Retention's second half deletes from `messages` alone and touched neither `files` nor `sessions`, so it moved no generation. The argument was that those rows answer nothing, which was true of retrieval and not of the engine: the typo repair's dictionary is a view over the FTS b-tree and listed them, so a drain running between two pages swapped the repair under a cursor that was still honoured, and a search that had answered stopped answering. The commit before this one fixes that at its source by counting live rows. It does not make the drain provably inert — the vocabulary still decides which candidates survive its scan limit, and reclaiming a term's last row moves where that limit cuts — so the fence is what covers the rest. A fourth trigger, on `messages`, with a `WHEN` clause that is the whole reason it is affordable: a replace and a `removeFile` delete a session's rows while its `sessions` row still stands, so neither fires, and both already bump through `files`. Only the drain deletes a row whose session is gone. The price is named rather than avoided: a cursor outstanding while a purge runs is now refused once per batch, which `SessionSearchCursorError` reports as `stale-generation` so a caller re-issues page one. The test that pinned the old contract is replaced by one for the new one, and by one proving a replace still does not fire it. |
||
|
|
b1514a1340 |
fix(ai-vault-search): repair a spelling inside the scope that will answer it
Typo repair read `messages_vocab` and probed `messages_fts` with no column filter, so tool output decided whether a conversation-scoped query was repaired, in both directions. A tool row carrying the misspelling made the query look correctly spelled and suppressed the repair; a tool row carrying a rare word became the suggestion, naming in `repairedTerms` a string from a column the scope will never show. Both reproduced against a control index that differs by exactly that one row. The vocabulary proposes and a scoped count disposes. fts5vocab is per table and cannot be column-filtered, so every decision that reaches the plan — already spelled right, eligible, and which of two equally close candidates wins — now comes from a `messages_fts MATCH` under the same filter retrieval uses, joined to `sessions`. That also takes the vocabulary's `doc` out of the ranking, which is the half of the drain defect that belongs here: `doc` counts rows whose session a purge has already cut loose, so reclaiming them changed which word a query was repaired to. Candidates are ordered by term now, because the ordering decides which of them survive the scan limit, and ties on similarity go to the more common word counted live rather than to the vocabulary's number. The cost is one bounded count per candidate examined, at most eight per prefix, and only for a term the scope has no posting for at all. |
||
|
|
adc4022658 |
refactor(ai-vault-search): answer the conversation scope with a column filter
PR 2 deleted `conversation_fts` on the strength of this PR's shoot-out, so the
scope is a column filter over the one FTS table now. `ftsTableFor` is gone; a
scope is a pair of `scopedExpression` and `scopedWeights`, and the table name
no longer travels through the engine, the snippet builder or a hit.
The filter is parenthesised, and that is the whole of it: `{cols}: (a AND b)`
binds both terms, while `{cols}: a AND b` binds only the first and searches
tool output for the rest. A test drives an AND whose second term lives only in
tool output through both scopes.
The snippet keeps one guard, not two. Its column list and its expression were
each hiding the other's mistakes — a tool-only row was unreachable through
either — so the list is the same four columns for every scope and the scoped
expression is what makes a conversation snippet impossible to draw out of tool
output. Dropping it now leaks that row, which a test catches.
One behaviour the deleted table did not have, pinned rather than wished away:
bm25 normalises by the whole row's length and has no per-column length, so two
rows with identical prose score differently when one also holds tool output.
The rowid set is unchanged; the order within it can move.
Re-measured on the shipping schema. The conversation scope is 1.2-1.4x faster
than `all` at every rung, and the index is 57 MB rather than about 150 MB at
93% tool output, because a tool row is now capped at 3,072 characters.
|
||
|
|
063283ce50 | test(ai-vault-search): pin that an append moves the generation a cursor is fenced by | ||
|
|
ce5c17e24c |
perf(ai-vault-search): price the second FTS table, and re-measure without warmup
Open decision 3. A column filter over `messages_fts` returns the identical rowid set as `conversation_fts` — checked here per query rather than assumed — so the table exists for latency alone. On a 105 MB corpus at both ends of the tool-output band, the column-filtered form costs 1.16-1.42x at p95, against a bar of 2x, so the recommendation is to delete it. The shoot-out writes its own corpus because the answer turns on the one property the shared generator fixes: how much of a transcript is tool output. Half the tokens in that output are words the conversation also uses, which is deliberately generous to the table under question. The doc records the number that argues the other way. PR 2 priced the table at about a quarter of the index on a corpus whose tool output is 56% of its message text; on a tool-heavy one it is 6.7-11%, because `messages_fts` grows with the tool text and the second table does not. Page warmup is not re-added. The measurement behind it was on a 4 GB index, removing it moves this corpus by less than the run-to-run spread, and a cancellable background pass needs a lifecycle a query library does not have. |
||
|
|
73cada03af |
refactor(ai-vault-search): read the tables the simplified index writes, and own the fence
PR 2 deleted the visibility views, the staging tables and the store's generation, so this reads `sessions` and `messages` directly and carries the three schema objects only a query needs — the vocabulary, the query log, and the triggers that move the generation — as its own extension over the store's schema. The fence is now three triggers on `files`, because every transaction the store opens that can change what a search returns writes that table and nothing else does; retention's orphan drain is the one write path that touches neither, and it is the one that must not bump. The triggers live in the file, so a writer in another process moves the generation without knowing a reader exists. A message row can now outlive its session row until the drain reaches it, so the snippet read joins `sessions` and the typo repair asks for a live posting instead of trusting the vocabulary's document count. The engine takes a connection rather than a store: PR 2's store keeps its connection private, and which process may open or rebuild the index file is PR 3b's decision, not a query engine's. |
||
|
|
1a3b48f130 |
fix(ai-vault-search): trim operator values, and report a query the engine had to cut
Three lows. `repo:" "` survived as a term and matched no label, silently emptying the list, which is the exact defect the empty-value drop exists to prevent wearing different clothes. Operator values are trimmed, and a whitespace-only one drops like an empty one. The substring matcher's copy of a free-text term is trimmed too, so `" "` reads as the empty term already does; the span kept for FTS is still the query verbatim. Two caps upstream of retrieval fired silently: the planner searches at most 48 terms, and the engine cuts the query at 512 characters. A 56-term query whose only match was the 56th came back with no hits and nothing truncated, which claims there is nothing to find. `truncated.query` now says when either fired, alongside the candidate and snippet flags that already did. Every cursor refusal now carries the generation the index is at, which the engine knows before it looks at the cursor, and the generation the cursor claimed wherever that survived parsing. The doc says exactly when each is present instead of leaving absence unexplained. |
||
|
|
791f1ed59d |
fix(ai-vault-search): restore the panel's reading of a quote that does not end a word
Unifying the two parsers changed panel behaviour on nine of twenty probed shapes, not the three previously pinned. Six of the nine were regressions, all from one rule: the shared parser refused a quoted span whose closing quote was not followed by a space, so `"a b"c` and `repo:"a"b` became single terms carrying their own quote characters, which match nothing. The rule was justified as what stops the apostrophes in `it's a repo:orca thing's` from swallowing the operator between them. It is not: a span only ever opens at a token start, and the quote in `it's` is not at one. Dropping the rule restores all six shapes to what the panel has always done and leaves that protection intact. Three changes remain and are kept because the old answer was worse in each: an operator with an empty quoted value is dropped rather than filtering on `""` and silently emptying the list, and a bare pair of quotes reads as an empty term rather than as the two characters. Each is pinned with a test that says which behaviour it is and why. |
||
|
|
ac8802d2e5 |
fix(ai-vault-search): stop reporting a search that gave up as a search that finished
The operator-only walk stops at a scan ceiling as well as at a full candidate set, and only the first of those reached the result. A query whose one match sat past the ceiling came back with no hits and truncated.candidates false, which is the engine claiming there is nothing to find when what happened is that it stopped looking. Retrieval now says why it stopped, because it is the only layer that knows, and the count it used to return could not distinguish the two cases. The cursor fence stays as it is: any published read moves the generation, so an outstanding cursor is refused, and that is what F11 asked for. What was wrong was the claim next to the row-delete skip that pagination stays usable through indexing. It does not, and the engine now says so. The rejection carries the generation the cursor was minted in and the one the index is at, so a caller can tell a moved index from a bad cursor and re-issue page one without showing anyone an error. The capability probe was nearly dead code, since every store opens through a function that rebuilds a stale file. It is not dead, because two handles can be open on one file, so the claim is corrected rather than the probe deleted. It now runs per search: a verdict cached in the constructor is wrong in both directions once another handle rebuilds the index. Also says why the row-delete loop may skip the bump: those messages keep batch_id NULL and stay in visible_messages, so what makes them unreachable is their session's tombstone, and the read ratchet is what keeps every reader joining the view that applies it. |
||
|
|
5a9557179f |
docs(ai-vault-search): re-measure the query benchmark after the operator change
repo: and path: moved out of SQL, so the operator-only row measures something different now and the note that it is a range seek was wrong. The rest of the table is re-measured on an idle machine: the previous run's p95 column was mostly contention, which is why conversation looked 3.4x faster at p95 rather than the 1.7x it actually is. Adds what this PR does not settle: which process may unlink and rebuild the index is PR 3b's, and PR 4 is only the first thing that makes reading it reachable. |
||
|
|
cc3de3cda5 |
fix(ai-vault-search): make repo: and path: mean one thing in the list and the index
The engine said `repo:` and `path:` a second time, in SQL, and SQL cannot say them. LIKE folds ASCII and nothing else, so `path:CAFÉ` missed `café`; the engine searched `cwd_key` while the panel searches the working directory and the transcript path, so `path:jsonl` matched every session in the panel and none in the index; and the engine compared one path segment where the panel compares the last two, so `repo:orca/session-search` missed. All four reproduce in both directions. So there is one definition now, not two that resemble each other. `matchesAiVaultQueryOperators` moves into the shared filter module beside the panel that already owned the semantics, the panel calls it, and the engine applies it over the rows it retrieved. SQL keeps only what it can express exactly: the `cwd_key` prefix range for `scopePaths`. `parseVaultQuery` now parses through `splitAiVaultSearchQuery`, so one parser decides what an operator is. Its existing tests pass unchanged; four degenerate shapes do answer differently and are pinned as decisions rather than left to be discovered. Two consequences worth stating. The operators are conjunctive now, because that is what the panel has always done, where the engine had been ORing within a key. And the operator-only page walks newest sessions in bounded pages applying the predicate, rather than taking one cut of the newest N and filtering it, which would have answered `repo:x` with nothing on a busy index. |
||
|
|
19a7c6a8fa |
fix(ai-vault-search): answer from a version-1 index, and keep a repaired literal whole
Two smaller findings. An engine can be handed a connection to a version-1 file that another handle is still answering from, and this PR is what first makes that reachable. It now probes once for the two tables version 2 added and names what it cannot serve on every result, instead of throwing at the first query that reaches for a vocabulary that is not there. The route ladder simply skips its repair rung. Which process may unlink and rebuild the index is PR 3b's decision and is not solved here. Typo repair re-planned the corrected query from scratch, and a corrected spelling can read as prose even when what was typed was a literal: `parseJsonn(the, data)` has the punctuation, `parsejson the data` does not, so the re-plan dropped `the` as a stop word. The repaired query searched for less than was asked and `repairedTerms` reported a body nobody typed. The re-plan is now told what the original decided, because a repair changes spellings, not the query's character. `repairedTerms` is documented as the whole body the repaired plan ran, with the index's own case-folded spelling for the terms it corrected. |
||
|
|
c9b1f95875 |
fix(ai-vault-search): fence the generation against writers this process cannot see
The generation was cached in memory and moved only on this store's own writes, and the bump itself was a read-then-write outside any transaction. Two handles on one file is the normal case once PR 3 lands: the indexer writes in the scanner child while an engine reads elsewhere. A reader would see that writer's deletions while its own generation stood still, honour a stale cursor, and skip a session; two writers could mint one generation for two snapshots. `generation` now reads `meta` on every call, and the bump is a single `ON CONFLICT DO UPDATE ... + 1` inside the transaction that makes the change. That closes the crash hole the bump-on-open existed for, so opening a store no longer invalidates anyone's cursor. Only a change to what a read returns counts. Retiring a path the index never held hides nothing, and the row deletes that drain a tombstone take away rows that were already invisible: bumping there would refuse a cursor every 256 rows and leave pagination unusable for as long as indexing ran. Also drops redaction from the query log, following PR 2's decision to store transcript content as written; the module it called no longer exists. |
||
|
|
b0971246db |
test(ai-vault-search): pin the two engine seams nothing was holding
The warm wiring and the query-length cap were both written and neither was observable. A spy pins that a search is what warms the store, and the cap is pinned through the one input the planner's own term limit does not already bound: a single enormous token, where the cut is what decides whether the term matches the indexed one at all. |
||
|
|
49f6014380 |
refactor(ai-vault-search): key a page once per search, and report reachable pages honestly
The page key was hashed twice per search, once to decode the incoming cursor and once to mint the outgoing one. The benchmark's reachable-page figure was clamped against a constant that could never bind; it is the limit over the page size and nothing else. |
||
|
|
fa8295be06 |
perf(ai-vault-search): measure what a query costs and what the candidate limit buys
Two corpora, because they answer different questions. The 10.5 MB / 40-session corpus is what the scope split costs a reader: conversation is about 1.6x faster at p50 and 3.4x at p95 than the full corpus, which is the argument for the second FTS table being the one a keystroke can afford. The candidate limit needs more sessions than the limit before it costs anything, so it is swept over 2,500 one-turn transcripts with every session matching. Limits are interleaved sample by sample: run back to back, the first configuration pays for every page the OS cache had not seen and the ordering alone moved p95 further than the limit did. The doc says plainly what these numbers do not cover. They are cost, not relevance; the MRR figures quoted beside the BM25 weights and the identifier shadow column come from a shoot-out over real transcripts and cannot be reproduced from this repository. |
||
|
|
cc9beab637 |
feat(ai-vault-search): rank, page and answer a session search
`SessionSearchEngine.search()` over the PR 2 store: route ladder (phrase, AND, typo repair, OR), BM25 weights per corpus, one hit per session, fork folding, and a page. - `scope` picks the corpus and the engine never second-guesses it. `conversation` is user and assistant turns; `all` adds tool output and the identifier shadow column. Switching corpus while typing is PR 7's policy; an engine that widened on a miss would make a result impossible to reproduce from its own request. - Pagination is an offset into one ranked list, fenced by the index generation and by a hash of everything that changes the ranking. A cursor from another generation or another query is refused with a typed error rather than silently re-run. Ranking breaks every tie by session id, because a cursor indexes into that order and retrieval does not promise one. - Snippets and source presence are paid for by the page, not the list. A snippet past the per-hit ceiling is cut on a code point, never between `[[` and its `]]`, and flagged on the hit. - Source presence is read from the `files` table. No stat on the query path, and no `missing`: only a proven deletion may claim one, and this read cannot prove it. - `SESSION_SEARCH_CANDIDATE_LIMIT_DEFAULT` is an option, not a constant, and the result says when the limit was what cut the answer. - The engine is the only caller of `store.warm()`, on first search. Library only: no IPC, no settings, no Electron, nothing constructs it in production. |
||
|
|
669ac965d5 |
feat(ai-vault-search): narrow a search the way the sidebar keys a folder
Every caller-supplied narrowing in one place, so retrieval, the operator-only page and the session load cannot drift apart: agents, an updated-at floor, the retention cutoff, scope paths, and the repo: / path: operators. Scope keying goes through PR 2's `cwdKey`, which is the sidebar's `folderGroupKey` without its prefix, rather than the original branch's second spelling. That drops the branch's WSL distro qualification, which PR 2 removed on purpose, and it makes the filesystem root a key that already ends in a separator, so the child-prefix range is built from the key rather than by appending one; `//` sorts below every real child and would scope the root to nothing. Engine types land here too, under src/main and not src/shared: nothing in this PR is a wire type, and PR 5 lifts what a caller may receive. |
||
|
|
67318a8e09 |
feat(ai-vault-search): plan a query the way the index tokenized it
The planner unfolds the tokenizer contract instead of asking SQLite: same boundaries as `unicode61 tokenchars '_.-/+'`, pinned against real fts5vocab output so a query can be planned without a round trip. It decides the route ladder's first rung (literal shape), strips stop words from prose but never from a literal, and fans an identifier out into its pieces for the OR fallback. Typo repair uses the index's own vocabulary as its dictionary, so it can never suggest a term the index does not hold. Its exact-match probe now joins `visible_messages`: `messages_vocab` is a view over the FTS b-tree and still lists staged and tombstoned terms, and the PR 2 read ratchet is right to demand the join. `splitAiVaultSearchQuery` is the index's reading of repo: / path:. It keeps operator case, which the panel folds and cwd_key must not; a census test pins the two parsers to the same answer about what is an operator until PR 7 moves the panel onto this one. |
||
|
|
949427f787 |
feat(ai-vault-search): give the index a generation a page cursor can be fenced by
A search page is a slice of one ranked list, so a cursor only means anything against the snapshot that produced it. The store now keeps a monotone generation in `meta`, bumped by every mutation that can change which rows a read returns, and starts a new one on open so a write whose bump never landed cannot leave a cursor pointing at content that is already gone. Schema version 2 adds the two tables the query layer needs: `messages_vocab` (fts5vocab over messages_fts, the typo repair's whole dictionary) and `search_log`. Both are pure additions and the index is a cache, so a version-1 file is dropped and rebuilt exactly as any other mismatch is. |
||
|
|
9a56797486 |
fix(mobile): surface host create warnings and terminal-create errors (#20125)
* fix(mobile): surface host create warnings and terminal-create errors
A workspace created from the phone could land on "No tabs in this session"
with a bare red "Failed to create terminal" and no way to tell why. Two
independent drops hid the host's own explanation:
- createWorktreeWithNameRetry returned only {worktreeId, name}, discarding
worktree.create's `warning`, and hostNewWorktreeSessionRoute built the
session route with only `name` + `created=1`. The session screen has always
had the banner (MobileSessionContentRow + createWarningState) -- only the
tasks create path ever fed it, so the New Workspace path could never report
a startup terminal that failed to spawn.
- handleCreateTerminal collapsed every failure to the literal
'Failed to create terminal', throwing away response.error.message.
Both now propagate, so the daemon's pty-allocation hint ("Your system cannot
allocate any more pty devices.") reaches the phone instead of dying in the
main process. Behaviour is otherwise unchanged: a blank warning is still
omitted from the route, and a host that gives no reason still reads
'Failed to create terminal'.
* test(mobile): refresh route parity baselines
---------
Co-authored-by: Merge Sim <sim@local>
|
||
|
|
ec9c3e0550 |
Fix remote hosted review browser routing (#20030)
* Fix remote hosted review browser routing * Fix remote review modifier hint * Address review feedback and fix routing test types * Avoid assuming active runtime owns workspace links * Align runtime routing regression expectation * Respect explicit local link ownership |
||
|
|
fb19c969a6 |
feat(native-chat): show when a Codex goal is set, changed, or cleared (#19923)
* feat(native-chat): show when a Codex goal is set, changed, or cleared Codex never emits the model's `create_goal` call as an item, so `thread/goal/updated` is the only truthful evidence that a goal exists. Both goal notifications were classified `status-chrome`, which meant no typed handler read them and no row was written -- the only thing reaching the reader was the model's own prose. That prose can be wrong: in a session where no goal was ever created the model still wrote "Goal created: ...". Classify both frames as timeline-substantive and give them a sentence, so the reader can tell a goal that exists from one the model merely claimed. Status is translated rather than echoed, and an unrecognised future status still reads as "Goal updated: <objective>" instead of the bare opcode. Codex re-sends the goal as its token and time counters climb, so rows are deduped on what a reader would notice -- objective, status and budget. A live session sent the same goal three times in one turn with only accounting moving. * fix(native-chat): write the goal signature separator as an escape, not a raw NUL A literal NUL byte in the source made git classify the file as binary, which hid its diff from review. The string built at runtime is unchanged. * fix(native-chat): make Codex goal rows retry-safe * fix(native-chat): preserve Codex goal identity on resume * fix(codex): ignore empty goal clear snapshots * fix(codex): preserve goal lifecycle across rewinds --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
08a24efaba | fix(push): preserve distinct Android alerts while offline (#20066) | ||
|
|
9aa0f7e77d | Update README downloads badge | ||
|
|
a0799d8f1c |
fix(terminal): move the recovery ledger onto the tab row and gate it on observed outcome (#20025)
* fix(terminal): move the recovery ledger onto the tab row and gate it on outcome The recovery budget lived in module-level Maps keyed by tabId. Anything keyed outside the row needs a release path, and that release fired on every remount-driven pane disposal, so each remount erased the budget it had just consumed (crash b5cfc6ca). Put the ledger on TerminalTab and write it in the same set() as the generation bump: reading the budget is now reading the tab, so releasing it independently has no expression. Counting was also the wrong control. Every remount mounts a pane that captures a FRESH recovery epoch, so the epoch check can never refuse its request — recovery re-requested the exact action that had just failed with no evidence anything changed. Gate on an observed outcome instead, reusing the direct-SSH pane retry vocabulary (success | failed | timed-out | superseded) and its settle call sites: an unsettled attempt blocks the next one, and a settled failure refuses the same reason until a new trigger arrives (generation move, or the user's Retry). The 3-per-5min cap stays as a breadcrumb-emitting backstop, not the control. viewMode now also lands on the row from the local toggles, mirroring how pin already does it, so the chat-ownership guard reads one index instead of OR-ing two. * fix(terminal): persist the row's viewMode and keep both chat-ownership reads The narrowed chat-ownership guard read a field the session schema strips: terminalTabSchema never declared viewMode, so the terminal row lost it on every load while the unified tab kept it. After a restart the row read undefined and recovery would remount a chat-owned tab's hidden surface — the race #19745's guard exists to prevent. Declare viewMode on terminalTabSchema so the row is durable, and keep the disjunction rather than replacing it. The schema cannot retroactively add the field to sessions already on disk, so the first load after upgrade still has it only on the unified tab; and for a safety check over two partly-redundant sources, a hole in either index should err toward declining a heal. Also cover three structural guards that no test was holding: both remote ledger-carry paths (terminal-build, remote-workspace-session-merge) and the only success settle in the state machine, including its placement past the failure branches. * fix(terminal): settle a fresh spawn's outcome and prove the ownership guard across a reload spawn-left-pane-unbound was the one recovery reason with no success settle: its remount heals by spawning, not reattaching, so it reached none of the reattach settle points and left the attempt 'pending' for the full 31s bound. A fresh spawn that binds a PTY now reports it, the dual of the unbound settle that already reported failure. Two tests outside src/ still called remountTerminalTabForRecovery by its old boolean contract and broke CI; both are updated to the admission result. Also strips the client-local recovery ledger at the remote-workspace projection boundary, in the type as well as the destructure, so a future producer cannot put another machine's Date.now() on the wire. * fix(terminal): resolve the pane's tab row once for both epochs after the main merge #20034 replaced connect-pane-pty's inline tab resolution with findTerminalTabForPane, and this branch had rewritten the line below it to read the recovery epoch off the row that block used to bind. The merge was textually clean and semantically broken: `terminalTab` no longer existed, so typecheck failed and every test that connects a pane threw ReferenceError. Resolve the row once through the new helper and feed both epochs from it, which keeps #20034's refactor and this branch's reason for reading the row here — a second lookup would put another tabsByWorktree scan on the connect path. captureTabRecoveryGeneration is narrowed to the one field it reads so the helper's record type can carry it. |
||
|
|
94c2f96ea4 |
perf(daemon): stop scanning every cell for OSC links that cannot exist (#20077)
collectHeadlessOscLinkRanges walks every cell of every row on each snapshot, and called xterm's getCell without the reuse argument its own docs recommend, so a link-free scrollback paid a CellData allocation per cell for a guaranteed empty result. Skip the scan when xterm holds no OSC 8 registration, and reuse one cell when it does. Measured over a 5000-row link-free buffer at 200 cols, same harness back to back, median of 25: 43.85ms -> 0.00ms. This is our bug, not xterm's: xterm already reuses cells in its own serializer and documents the getCell(x, cell) overload for exactly this. |
||
|
|
20c56249d5 |
fix(terminal): keep a deliberately slept workspace cold until it is woken (#20075)
* fix(terminal): keep a deliberately slept workspace cold until it is woken Sleeping a workspace kills its PTYs but keeps its panes mounted and keeps each tab's session id as a wake hint. Any later remount of those panes (recovery, parking, portals) reattached that dead id, and the daemon's create-or-attach spawned a fresh shell, so slept workspaces revived on their own (#10205). The existing sleep-intent marker now outlives teardown and gates the deferred connect itself, so both the reattach and fresh-spawn arms stay cold. It is released by activating the workspace, by any PTY binding to one of its tabs (CLI, automation, client wake), and by purge. A queued startup still connects. Reproduces the community root cause from gatsby74 in #13343; the regression e2e remounts a slept hidden pane and fails on main. Co-authored-by: gatsby74 <gatsby74@users.noreply.github.com> Co-authored-by: mmarabel <mmarabel@users.noreply.github.com> * fix(terminal): let a slept pane wait for its wake instead of latching cold A pane whose connect ran while its workspace was slept used to mark itself connected and stop; nothing re-armed it, so a wake that produced a live PTY before the user clicked (CLI create, background agent resume, split panes) left panes stranded. The connect now waits on the sleep marker and resumes when the marker clears, and a torn-down pane drops its listener. Tabs created with a live PTY clear the marker too, the sleep flow marks each workspace only when its own teardown starts, and purge forgets the marker without waking anything. * fix(terminal): wake a waiting pane once, in its remounted generation Activation clears the sleep marker after the set() that bumps dead tabs' generations, and the waiting pane only resumes its connect when its tab generation is still current. Otherwise the stale pane and its remounted successor both reattached the same session id on a deliberate wake. * fix(terminal): resolve the waiting pane's tab by either id and re-arm after wake The wake listener looked the tab up by the pane's render id, which can be a unified id whose terminal tab lives under entityId, so the generation check declined forever for those panes. Mount, fresh spawn, and the wake listener now share one live resolver. The wait flag resets when the listener fires so a second sleep can hold the pane again, listener dispatch is guarded, folder activation clears after its own set(), and the sleep flow re-asserts the marker after each teardown while releasing a workspace the user activated meanwhile. * fix(terminal): ignore PTY binds that land inside the sleep teardown window A spawn resolving while shutdown was still awaiting the host bound a PTY and cleared the marker, waking every waiting pane mid-sleep; re-marking afterwards could not un-connect them. The sleep flow now scopes each teardown so binds in that window are not wakes. The e2e asserts a deliberate wake yields exactly one PTY, and the dispose test proves the listener is gone. --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: mmarabel <mmarabel@users.noreply.github.com> |
||
|
|
22d12388a5 | fix(pi): load extension providers for source control generation (#20070) | ||
|
|
78e985cd99 |
fix(pi): claim the status pane when the inherited owner PID is dead (STA-5245) (#16631)
* fix(pi): claim the status pane when the inherited owner PID is dead (STA-5245) The managed pi/omp/prime-agent status extension suppressed itself whenever ORCA_PI_STATUS_OWNED held a PID other than its own, with no check that the owner still existed. A restart leaves the previous owner's PID in the inherited env, so every later load returned early and the pane stopped reporting status permanently. Probe the owner before suppressing. Only ESRCH proves it is gone; any other probe result keeps suppression so a live foreign owner still cannot double-report. This mirrors the tri-state in main/agent-hooks/managed-hook-owner-identity.ts, which the extension cannot import because it loads inside the pi/omp runtime with no Orca deps. Also extracts the generated-source test harness into its own module so the suite stays under the max-lines limit. * fix(pi): validate inherited status owner pid markers --------- Co-authored-by: Neil <neil@stably.ai> |
||
|
|
6bb2b0c6d7 |
test(runtime): capture real Antigravity transcripts — the detector is inverted on live output (#19983)
* test(runtime): capture real agent PTY transcripts before rewriting Antigravity readiness
Antigravity readiness has been written five times against a five-line screen
typed from memory. There is no Antigravity transcript in this repository, so
every attempt was a guess tested against another guess. This adds the recorder,
the protocol and the fixture-driven suite so the sixth attempt can be written
against evidence, and changes no detector logic.
- config/scripts/capture-agent-pty-transcript.mjs records a live agent session
through a real PTY, escapes and wrapping intact. Ctrl-] is consumed by the
recorder and never forwarded, which is the only way to end a capture while a
dialog still owns the screen.
- config/scripts/pty-transcript-secret-scan.mjs finds account identifiers and
credentials, redacts them with same-length placeholders so wrapping survives,
and recognises its own placeholders so a scrubbed file verifies clean.
- src/main/runtime/antigravity-readiness-transcripts.test.ts asserts a verdict
per transcript and skips by name until the transcripts land, with a
doc-coverage ratchet and a guard that a fixture contains escape bytes.
The escape-byte guard exists because the three cursor-agent fixtures carry a
comment claiming they were captured verbatim through Orca, yet contain zero ESC
bytes and zero carriage returns. That comment is corrected here to say what
those files are; the fixtures and the rules built on them are untouched.
* test(runtime): capture real Antigravity transcripts, and pin what they prove
`agy` 1.1.25 turned out to be installed, so the transcripts this scaffold was
built for now exist. Six are recorded from live sessions and committed; the
rest are named as skipped, because reaching them would mean signing the
operator out or deleting their config.
The captures invert the story. On real output the shipped detector refuses a
genuinely ready screen and accepts a live `/model` picker:
- Antigravity paints a block-glyph logo down the left, so the model row never
starts a line. `startsWith('gemini', trimmedStart)` cannot match a real ready
screen, on any account or model. Stripping the logo flips the same screen to
ready, which means a decorative glyph decides readiness today.
- The `/model` picker prints `Gemini 3.x Flash` one per line, at line start, and
a bare `>` composer sits earlier in the tail. Both halves of the rule are
satisfied while a dialog owns the screen.
- For an API-key user the identity row reads `Gemini API key` — no `@`, no
domain — and `AGY_CLI_HIDE_ACCOUNT_INFO=1` removes the row entirely. The
account-row requirement of attempts 4 and 5 can never pass for those users.
- The banner is printed once and never reprinted after a dialog is dismissed, so
`headerIndex` cannot be the ordering anchor.
Four suite cases are pinned as KNOWN DEFECT: they assert what the detector does
so CI stays honest instead of permanently red, and flip to failing the moment
someone fixes it. No detector logic changed.
The recorder gains `--send "<ms>:<text>"` because a dialog capture has to be
driven and an unattended run has no TTY, and the scrub scanner gains a UUID rule
because agy prints a resumable conversation id on exit.
* test(runtime): capture agy mid-turn, and make the scan file reviewable
Answers the busy-frame question a P1 review raised against attempt six, with
two new captures from a live turn.
At the frame level the review is right: a busy frame parks the caret with the
same bytes as an idle one, `CR ESC[2A ESC[2C`, and the only differing row —
`esc to cancel` versus `? for shortcuts` — is erased by that park.
At the retained-tail level it does not reproduce. Each spinner tick is its own
repaint with its own `CR ESC[2A`, two rows higher than the frame's, which
splices the composer away: a live turn's tail ends on `⣟ Generating...`, with
no bare caret to match. A constructed input that keeps the park and edits only
the status text is not faithful, because a live turn has a spinner row
repainting below the composer.
The residual is the gap between a frame park and the next tick, where the tail
does end on the bare caret. Quiescence-gated paths are safe there because ticks
keep arriving; text-only paths are not, and for those the capture supports one
clause: a braille glyph on the last visible line means working. That predicate
already exists here for cursor-agent and should be reused, scoped to the last
line — a first-run transcript prints `⠾ Signing in...` during startup.
Also in this commit, from the same review:
- pty-transcript-secret-scan.mjs held raw 0x00-0x1f bytes in a character class,
so the one file gating real PTY data into history was binary to git and
unreviewable in a diff. It now tests codepoints, which the formatter cannot
fold back into control bytes.
- Pin `src/main/runtime/__fixtures__/*.txt` as -text. A Windows checkout would
otherwise normalise line endings and rewrite the CR bytes that make these
files evidence.
The recorder now stops appending at the stop moment rather than through
shutdown: an agent repaints an idle frame on its way out, which was overwriting
the mid-turn state the capture existed to record.
* test(tooling): allowlist the transcript scan test in the batch-shim ratchet
pty-transcript-secret-scan.test.mjs asserts that the capture recorder routes
an 'agy.cmd' shim through cmd.exe, so the shim literal it names is the
assertion, not a spawn. Fits the existing assert-on-shim-files category.
|
||
|
|
2167cd2994 |
test(monaco): drive the real Monarch tokenizer instead of walking rule tables (#19981)
* test(monaco): drive the real Monarch tokenizer instead of walking rule tables
* fix(monaco): stop a truncated JSONL record poisoning every record after it
`@string` was pushed unconditionally, and Monarch state survives the line
break, so one truncated record left every later record inside the string state
— the whole rest of the file rendered as a single string. A truncated record is
a normal way for a .jsonl log to end.
Gate the push on a lookahead proving the closing quote is on this line, and
consume an unterminated remainder as `string.invalid` without pushing anything.
Measured against the real tokenizer over realistic JSONL. Two alternatives were
rejected: collapsing the string into one regex fixes the poisoning but loses
every `string.escape` token on well-formed records, and `includeLF` does not
fix the primary case at all (the `[^"\\]+` content rule swallows the newline
before an EOL rule can match). This candidate's token stream is byte-identical
to the previous grammar on every well-formed record — the existing inline
snapshot did not move — and is at or faster than it on 20k pathological lines.
* test(monaco): pin monaco's own mdx grammar as the embed-recursion regression
Upstream's shipped mdx grammar enters a `js` embed on every `{` and pops
on `}` with no budget, so it reproduces the unbounded embed-entry
recursion exactly — evidence the shape is monaco's, not something Orca's
grammars invented, and a tripwire for a monaco upgrade that changes it.
The same file proves an Orca grammar stays inside the budget under the
identical line. Measured here: 500 interpolations -> 500 frames, no error;
3000 chars -> RangeError at 945 frames. The ceiling is runtime-dependent,
so the test asserts the failure, not the number.
Adds a regression for astro's `^`-anchored frontmatter pop rule
(monaco-editor#1127) so an indented or trailing `---` cannot close the
fence early, de-duplicates the line-cap constant onto the budget module's
`MAX_TOKENIZATION_LINE_LENGTH`, and renames the recursion suite: embeds
cannot nest, so "embedded recursion depth" described the wrong thing.
|
||
|
|
5fa62feda7 |
perf(terminal): mount only the visible pane on a worktree switch (#20034)
* perf(terminal): mount only the visible pane on a worktree switch Activating a worktree mounted a TerminalPane for every tab it holds, not just the one on screen. Cold-activation deferral existed for this but engaged only past four deferrable hidden tabs, which exempted the 2-5 tab worktrees that make up almost every real switch. Deferral now engages for any deferrable hidden tab, and the siblings it skips are admitted one per idle frame after the reveal, capped at the population the old threshold would have mounted eagerly. Steady-state pane, WebGL-context and heap population are therefore unchanged; only the frame the mounts land on moved. * fix(terminal): judge admission eligibility on the largest deferred set seen Review found the launch worktree never warms up: it is restored active before hydration opens the startup gate, so admission read an empty deferred set, cached ineligible, and never recomputed once the real plan landed. Judge on the high-water mark instead - an over-cap worktree still stays ineligible as its set drains, but a later plan is seen. Also from review: the e2e WebGL counter read getPanes(), which returns a public projection with no webglAddon field, so it was always 0; read getRenderingDiagnostics() instead. Filler worktrees now clean up on failure (testRepoPath is worker-scoped), and the restore metric is named for what it measures rather than implying a pixel assertion. * test(e2e): wait for the reveal to restore, and scope the latency budget off CI CI failed with 'revealed terminal never restored its content': the harness sampled a fixed 4s window, which a shared runner can outlast, so a slow restore was recorded as no restore. Poll for the restore instead. Also stop asserting a latency budget on CI. Shared runners cannot hold a threshold; the structural invariants (one pane mounted by the switch, warm set restored) are exact and stay asserted everywhere. |
||
|
|
fab78c7669 |
fix(native-chat): show one live-turn indicator, and make Thinking mean reasoning (#19977)
* native-chat: render one indicator row for the live desktop turn The turn-timing row and the spinner+activity line were two rows saying "Working" at once. A settled turn keeps its own row; the live turn now has only the spinner row, labelled provider activity -> Thinking -> Working for N through the shared resolver. Reasoning is the turn's content, so it no longer becomes the activity label, and "Thinking" now means the turn is reasoning right now rather than that it has produced no output yet. * mobile: give the live turn row a spinner and the shared indicator label Mobile's per-turn row is already the only live indicator on the structured lane, but it pulsed a bare word and never showed what the provider said it was doing. It now renders a spinner beside the same resolved label desktop uses, and reads reasoning from the journal instead of inferring it from missing output. The bridge lane's four prompt/interrupt write seams move to one module so the controller stays under its line cap. * codex: mark streamed reasoning as reasoning too, and pin the provider markers The settled reasoning item carried the marker but the streaming one did not, so a live Codex turn - the only time the indicator is on screen - never read as reasoning. Both paths now stamp it; a plan document keeps its own presentation and must never read as reasoning. * fix(native-chat): tighten live turn reasoning state --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
3b82d8de64 |
fix(runtime): let connections own host status recovery (#20003)
* fix(runtime): let connections own host status recovery Verify runtime status after authenticated connection recovery and publish ordered snapshots to desktop and browser viewers. Consolidate failed-status retries in the connection owner and remove renderer retry/diagnostics merging. Adapt sidebar host-state derivation and regression coverage from Omar Shahine's original fix in https://github.com/stablyai/orca/pull/19163. Co-authored-by: Omar Shahine <10343873+omarshahine@users.noreply.github.com> * fix(runtime): show blocked hosts honestly and remove obsolete status options * fix(runtime): preserve timeout guidance and update IPC test fixtures * fix(runtime): preserve status evidence and address review gaps * test(sidebar): assert workspace host icons dimming and recovery tooltips * fix(palette): require available hosts before adding implicit badges * fix: retain disconnected host snapshots for new renderers --------- Co-authored-by: Omar Shahine <10343873+omarshahine@users.noreply.github.com> |
||
|
|
b6e4457552 |
fix(worktrees): close an idle structured chat on delete instead of refusing (#19762)
* fix(worktrees): close an idle structured chat on delete instead of refusing `worktree rm` refused whenever any structured chat session was attached to the workspace, so an idle Codex/Claude chat that had already answered was harder to delete than a terminal actively running the same agent. The PTY sweep stops every terminal it owns and refuses only for the ones whose exit it could not verify. The structured sweep refused on `live` alone and never attempted the close, which ran only under force. `live` is lease state — a provider child is attached — not work in flight, so it was never the right proxy for "you would lose something". Close first, refuse only on what did not settle. The refusal now means the same thing the unverified-PTY one does, so the toast takes that wording. * fix(worktrees): fence, bound and word the structured-session sweep Review follow-ups on the close-first structured sweep. The close-first direction is unchanged; four things it got wrong are not. Host fence. `listLiveStructuredSessionsForWorktree` matched on `location.workspaceId` alone, and a `repoId::path` id names a DIFFERENT workspace on every host (STA-4343). Once the sweep started closing rather than refusing, deleting a local workspace could close a live chat on an SSH or paired-runtime copy of the same id. It now takes the same two host fields the PTY sweeps already fence on, compared against the session's own `location.executionHostId`; neither field set means this machine. Shared budget. The close ran to completion before the first PTY sweep was constructed, and it is serial with a provider round trip per session — so a slow one spent the whole budget and the sweeps then rejected with a timeout for a stop they never attempted. It is now issued first but joined before the verdict, so the agent plane is still asked ahead of the terminal plane while the two share the clock. Timeout wording. The close raced the deadline fail-closed, and that sentinel carries the PTY timeout prefix, which the classifier reads first — so a wedged session close refused in terminal wording and refused identically again under the Force Delete meant to clear it (#11960). A close that ran out of time is now a session the removal could not confirm closed, which is what the refusal already words. Tracked, so a forced removal still waits out the abandoned-sweep grace before deleting files. Verdict fidelity. `closeStructuredAgentSessionChild` re-observes after the close, and that verdict was being discarded — so a session Orca watched stay attached and one it merely could not reach produced the same message, while the toast asserted "could not confirm" for both. `removal.ts` documents flattening those two as the thing not to do. The unclosed sessions now carry their post-close status, the detail uses the shared `still live:` marker, and the toast branches on it like the PTY pair above it. Also: the close takes the enumerated list instead of re-deriving it, so it no longer runs every liveness observation twice or names a session it never touched; and both teardown log lines count structured closes, since closing a chat is now an ordinary outcome of this verb. * fix(worktrees): keep a proven-exited session from refusing removal The structured sweep re-observes after a close that reported `stopped: false`, but folded a proven `exited` into `unverifiable` — so a close that threw past its own observation, or one whose death evidence landed a beat later, refused a delete over a child that is demonstrably gone. That is the defect this sweep exists to remove, and the PTY gate it mirrors never refuses on a proven exit. Take the proof, and run the tab retirement the close skipped when it gave up: a chat tab left behind re-attaches a released session pointing at a workspace that is about to be deleted. * fix(worktrees): name every unclosed structured session, not just the live ones The refusal named only the proven-live subset when any session was live, so a sweep that left one attached and two unconfirmed told the user "1 agent session (claude)" while three were about to be discarded — and dropped the providers of the ones it hid. The PTY sibling may drop everything outside its live list because a fresh inventory PROVED those exited; nothing proves that here, so both groups are counted. The `still live:` marker still leads, so the delete toast keeps showing the stronger warning. Also carries the structured close count through the forced-removal early return: that path skips the per-PTY verdict, not the sweep that already ended a user's chats, so the removal log claimed `structured=0` for chats it had just closed. * fix(worktrees): stop the forced-removal warn asserting a verdict it does not have The structured sweep splits its post-close verdict in two on purpose: "we watched it stay attached" and "we could not confirm it closed" are different things to waive, and `removal.ts` keeps a marker and a matcher together so the delete toast can tell them apart. The force-path warn then appended "still attached" to whichever verdict it got, so a removal forced over a close that merely ran out of time logged that Orca had seen the session running. That line is the only record a forced removal leaves of a child left pointing at a deleted `cwd`, so it is the one place the two must not be flattened. Carry the verdict verbatim, the way the unstopped-PTY warn above already does. * fix(worktrees): report the closes that landed when the sweep budget expires The structured close loop is serial, so the shared sweep budget can expire part-way through it. The timeout fallback was assembled by the caller and could only name the whole list: sessions this removal had already closed were reported as unclosed, named in the refusal the user reads, and logged as `structured=0`. The loop now records progress into a structure the timeout path reads, so both the refusal and the count say only what was observed. A session with no recorded outcome reports `unverifiable` — the same verdict as an attempted close that stayed unproven, because "never asked" and "asked, unconfirmed" are both exactly "not observed exited", and neither may claim `live`. The loop also checks the deadline before each close, so one slow provider round trip no longer starves every session behind it. It stops ISSUING closes; an in-flight one is left to finish, since nothing here can cancel a round trip. The structured host fence now reuses the PTY fence's own type instead of a look-alike that read `undefined` as local while the other read it as match-all, with both claiming the same precedence. `null` means this machine on both sides; ABSENT stays narrowed to local here, documented and pinned, because a single-host-id comparison cannot express match-all. Also pins a tradeoff that was accepted rather than wanted: the PTY sweeps run concurrently with the structured close, so a removal that refuses over a stuck session has already killed that workspace's terminals. * fix(worktrees): put the chat tab back when a structured close does not land `closeStructuredAgentSessionChild` hides the session's chat tab before it issues the close, so every failure past that point left a refused delete having still taken the tab out of the durable restore index. The conversation survived under `userData`, but nothing brought the tab back at the next launch. Both failure shapes now roll the hide back: `host.close` throwing, and the post-close observation coming back not-`exited`. The restore is gated on the visibility read taken BEFORE the hide, so it never publishes a tab for a session that was already hidden, and on a fresh observation, so it never resurrects one for a child a throwing close still took with it — which is what the worktree sweep reads when it counts such a session closed. It cannot throw out of the function, so the caller's original reason is still what the user is asked to act on. * fix(worktrees): keep the chat-tab rollback out of removals that delete the workspace The rollback added for a refused close ran on every unproven close, including the two shapes of removal that cannot refuse. Force Delete warns and deletes the checkout; a folder-workspace removal never refuses at all. Putting the tab back on those paths leaves a durable restore-index entry for a workspace that is then gone, and the chat republishes at the next launch pointing at it — the outcome this sweep exists to remove. The close now takes `restoreTabOnUnprovenClose`, on by default so `worker-stop` and `worker-release` keep the rollback, and the teardown sweep passes it only when the removal can still refuse. Second hole, same chain: `host.close` can return before the child's exit is recorded, so the close's own observation reads unverifiable and restores the tab, while the sweep's re-read one store write later proves the exit and counts the session closed. The two observations straddle that write and disagree. The sweep now re-drops the tab reference when it takes that proof, and the comment claiming the re-observation alone covers this is corrected. * fix(native-chat): stop a closing chat reading as a conversation that would not load Deleting a workspace now closes the structured chats inside it, and the chat pane outlives that close by a few frames. Every read it makes in that window — `agentSession.history` on refresh, `agentSession.subscribe` on reconnect — resolves through the host's `requireSession`, which refuses with `agent_session_ownership_unknown` for a session it no longer holds. The pane turned that into its terminal error surface, so an ordinary delete flashed `Could not load conversation` over the transcript before the tab retired. That code, raised by a READ, never means the transcript could not be read. It means this host has no session object by that id: one it has just closed, or one it has not attached yet, since the surface's hold is what attaches a session at all. Both windows end on their own. The genuinely latched lease — Orca cannot prove the previous owner exited — reaches the client through the acquisition path instead, so narrowing on the code costs a read no real diagnosis. So the read transport classifies before it reports: an unattached refusal stays on the reconnect loop it is already the subject of, and the pane keeps the transcript it has. It is a window, not a mute. A read still refusing that way past the grace is no longer transitional, and the pane is owed the failure rather than a spinner that never resolves. Every other failure still surfaces immediately, unchanged. The refusal code now has one definition, shared by the host that raises it and the client that narrows on it, so the two cannot drift into a red error nobody meant. Deliberately NOT changed: the order of teardown. The tab is retired after the close proves, not before it, because a close that does not settle has to put the user's chat tab back — the rollback this PR already establishes. Retiring the pane first would unmount it ahead of a close that may be refused, so the pane instead treats a session that has gone as a neutral terminal state. Mobile's structured chat reaches the same reducer but has no reconnect loop, and its hold refusal is what carries the diagnosis there, so the grace does not transfer; it keeps reporting as before. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
b3e0a33fa4 |
fix(runtime): agent-neutral wait-blocked reasons (#19749)
* fix(runtime): agent-neutral wait-blocked reasons and non-Gemini Antigravity readiness Reported by a user via the in-app help menu (report "not captured", 1.4.198). The trust/interactive/update/cwd prompt matchers are agent-agnostic - they match on dialog wording and never inspect the pane's agent - yet emitted hardcoded codex-* reasons. Those reached users verbatim in worker receipts (local-worker-start, federation), two automation surfaces, and raw CLI output, so an Antigravity user was told they had a Codex problem. findAntigravityReadyPromptIndex also required the model line to start with the literal "gemini". Antigravity CLI is not Gemini-only, so a non-Gemini session never registered as ready, stale trust text was never superseded, and the pane stayed blocked - which is why dispatch --inject answered agent_prompt_blocked. Add agent-neutral reasons additively (codex-* members kept on the wire per docs/reference/remote-wire-compatibility.md, with a legacy alias for older hosts) and decide Antigravity readiness structurally: header, then model/account rows, then the prompt caret. codex-model-migration-prompt and codex-hooks-review-prompt stay Codex-named - both key on Codex's own wording. * fix(runtime): finish the agent-neutral rename, revert the Antigravity readiness rewrite Review follow-up on this branch. Splits the two halves of the original commit: the reason rename lands, the Antigravity readiness detector goes back to merge-base until someone captures a real transcript. Rename half: - 'hooks need review' + 'press enter to confirm' inspects no agent, so it now publishes agent-hooks-review-prompt. That was the last agent-agnostic codex-* emission left, and it is the one the original report was about: a Claude Code user hitting a hooks dialog still read "codex-hooks-review-prompt". - The legacy alias is applied at all three surfaces that render a raw reason, not just the CLI. describeTerminalWaitBlockedReason() is the single formatter; the worker and federation "Agent startup blocked:" receipts use it too. Kept one-directional: nothing consumes agent-* -> codex-*, since an old client renders with its own shipped code. - Restores the compat note deleted at the permission-choices site. The Rule 1 citation is correct - remote-wire-compatibility.md names this enum by name. Antigravity half, reverted: findAntigravityReadyPromptIndex goes back to merge-base (header + a 'gemini' model line + a lone '>' caret) and antigravity-ready-prompt-index.ts is removed. Executing both builds against constructed tails, the rewrite read a live startup dialog as ready. Adding the account row from this repo's own ready-screen fixture to five silent startup dialogs (sign-in, model picker, theme picker, privacy notice, update banner) flipped all five from unready to ready; so did any narration line containing an email address, with no account row at all. Readiness is what gates typing the task prompt into the pane, so that path types a task prompt into a live authentication dialog. Merge-base returns unready for all ten. The rewrite also did not reliably fix the wedge it targeted: with no account row and a non-Gemini model - a personal or API-key user - it still returns unready. No real Antigravity transcript exists in this repo. The cursor-agent rules are derived from captures under src/main/runtime/__fixtures__; Antigravity has no equivalent, and every attempt so far has been tuned against a hand-written 5-line fixture. A false negative (the agent waits) is safer than a false positive (we type into an auth dialog), so this ships the known behaviour. Reverting restores a pre-existing gap, not a regression: a non-Gemini Antigravity session wedges on merge-base too. Closing it needs a captured ready screen and a captured dismissed-dialog screen, for a personal/API-key account as well as a Business one. Tests: - Ten ratchet fixtures pin the shapes any replacement detector must refuse - the five silent dialogs with an account row, and each with a narrated email. All ten fail against the reverted rewrite. - Vacuous tests rewritten so they fail without the code they cover: the CLI alias tests asserted only the absence of a suffix, and the worker receipt test asserted the raw token. Tests that are characterization rather than a guard now say so on the line above. --------- Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
26db907895 |
fix(monaco): bound embedded-language recursion in svelte/astro/vue grammars (#19748)
* fix(monaco): bound embedded-language recursion in svelte/astro/vue grammars
Monarch's _nestedTokenize and _myTokenize tail-call each other on every
mid-line embed entry, and V8 has no TCO, so JS stack depth grew one level
per <script>/<style>/<!-- -->/{expr} on a line - bounded only by line
length. One 17,000-char line of '<script></script>' overflows svelte at
depth 997; 19,603 chars of '<!---->' overflows astro at 1748. Both are
under Monaco's own 20,000 maxTokenizationLineLength, so it was no defence.
The same recursion rescans the line remainder per level (quadratic),
matching the 38s synchronous stall before report 25d10fa1's
STATUS_STACK_OVERFLOW on a flat, healthy heap.
Add a shared embed-entry budget: enter an embed only while <=512 chars
remain. Each entry consumes a character, so depth is bounded by
construction; over-budget remainders continue on parallel non-embedded
states that keep tag-level colouring.
Also fix the zero-width nextEmbedded rules that dropped their embed
(token must be '@rematch'), the source of the "cannot pop embedded
language if not inside one" breadcrumbs, and rework vue's expression exit.
Fixing that drop without the budget would have armed the overflow in vue.
* fix(monaco): re-embed script and style bodies after an over-budget opening tag
`scriptBodyPlain` / `styleBodyPlain` were the only over-budget mirror states
without a re-entry rule, and they dropped `$S2` as well. A `<script>` or
`<style>` opening tag carrying more than 512 trailing characters therefore left
the whole block unhighlighted until its closing tag, however short the following
lines were. Carry the language through and re-enter the embed as soon as the
rest of the line fits, matching the markup and expression mirror states.
Also extends the recursion ramp so the densest embed shape (`{a}` / `{{a}}`) is
driven at Monaco's line cap (19_800 / 19_528 chars) instead of stopping at
7_500, and renames the inverted private `restOfLineTooLong` constant.
The budget stays at 512: an A/B of the real tokenizer at 512 vs 256 over
realistic SFCs differs on 8 lines, all of them 256 losing the html or
typescript embed on ordinary shapes such as a ~430-character Tailwind class
attribute.
* fix(monaco): pin the tokenization line cap and correct the budget's claims
Sets `maxTokenizationLineLength` explicitly instead of inheriting it. It is
an `IGlobalEditorOptions` value, so the one file-editor site pins it for
diff and Peek surfaces too. Defense-in-depth only — the comment says
plainly that it does NOT guard the embed recursion, which overflowed at
~17_000 chars, under this cap.
Corrects two overstatements in the budget module's own comments: Monarch
refuses to nest embeds, so the counts are sequential enter/exit
transitions (stack frames), not nesting depth; and the `RangeError` is
caught per line by Monaco's `safeTokenize`, so what is demonstrated is a
line that silently loses highlighting, not a dead renderer.
Notes monaco-editor#1127 at astro's `^`-anchored frontmatter pop rule: the
two-caret symptom is fixed in 0.55.1, but `^` in a pop rule is still
measured from where the embed was entered, not from line start.
---------
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <neil@stably.ai>
|
||
|
|
7d367b2aa4 |
fix(terminal): stop the recovery budget erasing itself during a remount (#19745)
* fix(terminal): stop the recovery budget erasing itself during a remount A recovery remount disposes the pane's xterm, and the disposal handler released the tab's recovery budget whenever getTab said the tab was gone. getTab reads unifiedTabsByWorktree while remountTerminalTabForRecovery reads and mutates tabsByWorktree; on the direct-SSH path the two indices diverge, so every successful remount deleted the timestamp it had just written. The cap never engaged and the pane remounted at render speed. Report b5cfc6ca (1.4.198, Windows): 8878 remounts across 8 tabs in 122s, all reason=reattach-unverifiable, against a cap of 3 per tab per 5 min, ending in a Skia bitmap allocation abort. Gate the release on the same index that governs remounting, via a shared locateTerminalTabForRecovery so the two cannot drift apart again. * refactor(terminal): resolve a recovery tab through one tabsByWorktree scan Collapse the recovery lookup onto a single primitive, locateTerminalTab, and express the already-exported isTerminalTabPresent in terms of it. The budget release now calls isTerminalTabPresent directly, so the store drops the hasTerminalTabForRecovery action added alongside it. The native-chat ownership guard also consults tabsByWorktree: the unified tab index owns viewMode, but it can transiently drop a row the remount index still holds, and that hole used to read as "not chat-owned" and remount a chat-owned tab. Drops the storm control case that passed with and without the fix, and pins the surviving drift case to the exact remount count the cooldown produces. --------- Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local> Co-authored-by: Neil <neil@stably.ai> |