mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 08:02:12 +00:00
2222e5475480bb808cfcd39e2ac3204c75bd3dc1
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
97b71c2285 |
refactor(usage): split AI-usage scanners and stores under the max-lines budget (#14668)
The three usage scanners and their stores, plus the renderer usage-overview model, each carried a file-level `eslint-disable max-lines` and had grown to 338-769 counted lines against a 300-line budget. AGENTS.md calls for splitting rather than suppressing, and config/max-lines-baseline.txt is a shrink-only ratchet, so this removes all seven suppressions and prunes their entries (341 -> 334). Each file is cut along the seams it already had -- and that several of the suppression comments named out loud: filesystem discovery / record parsing / attribution / aggregation for the scanners, and pricing policy / scope filters / rollups / session rows / automation attribution for the stores. Pure move, no behavior change. Code is relocated verbatim; the only edits are import plumbing and, where a private class method became a free function, the mechanical `this.state` -> `state` parameter threading. Every converted call site passes `this.state` at call time and the automation path takes a live `getState: () => this.state` getter, so no state is snapshotted. No barrel exports: each new module owns real logic and importers point at the owner. Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green (remaining failures are pre-existing load flakes in untouched files, each green when re-run serially), no import cycles among the 64 affected modules, and a statement-level diff of every split confirms the moves are verbatim. |
||
|
|
65e2b5b598 | refactor(usage): share provider store lifecycle (#13558) | ||
|
|
07bd574294 |
refactor(usage): share the Codex/OpenCode scan fold behind a provider contract (#12082)
* refactor(usage): share the session/daily fold between Codex and OpenCode The Codex and OpenCode scanners each carried their own byte-identical copy of the ~325-line aggregation pipeline (createEmptySession, the three breakdown folds, finalizeSessions, mergeSessions, mergeDailyAggregates). Two copies means a token-accounting fix — a bucket that double-counts, a merge that drops a breakdown row — lands in one provider and silently not the other. The copies had already started to drift in comments only; the next drift would have been in arithmetic. The providers differ in exactly one dimension: the extra metric folded alongside the token counters (Codex `hasInferredPricing`, OpenCode `estimatedCostUsd`). That is now injected as an empty/fromEvent/fold triple, so the shared code stays generic without collapsing the two record schemas into a nullable union. The clone strategy stays per-provider (`cloneSessionForMerge` vs `structuredClone`) rather than being unified on the assumption that the difference is accidental. `usage-provider-contract.ts` is the seam a plugin-contributed usage source will implement. It is deliberately generic over each provider's record types: Claude bills per turn while Codex/OpenCode bill per event, and `cachedInput` is a subset of `input` for the latter but a peer bucket for Claude, so a single normalized record would push nullable handling onto every consumer. No behavior change. Emitted objects are byte-identical, including key insertion order — verified by diffing JSON.stringify of the scan output before and after across mixed models, mixed locations, an inferred-pricing flip, and null vs non-null cost. Persisted field names and schemaVersion are untouched, so caches do not invalidate. * refactor(usage): make the provider contract load-bearing and dedupe worktree refs Follow-up to the aggregation extraction, addressing three review points. `UsageProvider`/`UsageScanResult` were declaration-only, which is the same speculative-interface problem #12077 just deleted 8,900 lines of. They are now implemented by both real providers via `satisfies`, so the seam is typechecked against actual scan functions rather than asserted. The blocker was that codex returns `processedFiles` and opencode returns `processedDatabases`; rather than rename persisted-adjacent fields, the source key is a type parameter, so each provider keeps its own on-disk name and the contract still binds. Verified the constraint bites: swapping the key to 'processedSources' fails typecheck. `schemaVersion` is part of provider identity in the contract, so each provider's SCHEMA_VERSION constant (with its cache-invalidation rationale) moves into the provider module and the store imports it. Values are unchanged (codex 5, opencode 2) and the stores compare them exactly as before, so no cache invalidates. This also keeps store -> provider -> scanner acyclic. `UsageWorktreeRef` collided with the existing export in usage-worktree-metadata (3 fields, no repoId). Two different exported types under one name in src/main is worse than the duplication being removed, so the scan-input type is now `UsageScanWorktreeRef`; usage-worktree-metadata is untouched. `createWorktreeRefs` was triplicated. Codex, OpenCode, and Claude copies are byte-identical apart from the return type name (verified by diff), and all three ref types have the same four fields, so one shared copy replaces all three. This is the only change to claude-usage/. No behavior change: same functions, same arguments, same call order. The store tests' `./scanner` mock still intercepts scanning because the provider captures the mocked binding; their now-inert `createWorktreeRefs` mock key is dropped so it does not read as still mocking something. |
||
|
|
cbe8635f46 |
fix(worktrees): prevent deletion from blocking Orca (#11233)
* fix(worktrees): prevent deletion from blocking Orca * test(worktrees): loosen async history-delete event-loop bound for CI The main-thread safety check failed on a loaded runner when a single timer gap hit ~48ms under the prior 30ms threshold. Keep the bound well below a recursive sync-rm stall without treating CI jitter as a block. * test(worktrees): measure history-delete critical path, not timer gaps setInterval gaps during async rm of thousands of files still flake under CI scheduling. deleteWorktreeHistoryDir is sync and must only rename, so assert that critical-path wall time stays well below a recursive walk. * fix(worktrees): prevent deletion from blocking Orca Add timeout-based draining of watcher closes so SSH round-trip delays don't indefinitely block the worktree removal path. Also: order durable temp-file sweeps ahead of writes to reclaim orphans before accumulation, skip own-process temps to avoid deleting live writes, swallow persistence errors so disk failures don't cascade to query callers, and measure history-deletion progress by loop turns rather than timer gaps to detect blocking on CI runners. * fix(worktrees): prevent deletion from blocking Orca Worktree deletion can now proceed even if filesystem watchers or history cleanup operations hang, preventing Orca from freezing. Changes: - Fence install slots with tokens instead of counters so removals can abandon wedged installs without corrupting later removals - Timeout-bound watcher unsubscribe operations with a shared drain budget - Move JSON serialization of large usage caches from queue-time to write-time to avoid blocking main thread - Async tombstone + schedule history tree deletion instead of blocking recursive rmSync during GC, preventing main-thread stalls ~10s after startup * Extract usage cache writer into reusable durable snapshot class Consolidates serialized durable-write and generation-veto logic from three usage stores into UsageCacheSnapshotWriter. Eliminates duplication, centralizes multi-MB JSON serialization on the main thread via write-queue serialization, and vetoes superseded snapshots to avoid wasted rewrites. * fix(worktrees): prevent deletion from blocking Orca Worktree deletion used to recursively delete large session trees (hundreds of MB) on the critical path, stalling the event loop. Instead, rename trees into a `.pending-delete` tombstone queue and reclaim them asynchronously off the removal's critical path. Extracted host tree removal into a reusable helper (`removeHostTree`) that centralizes Windows retry logic. Added usage-cache flush on quit to prevent data loss when scans complete right before shutdown. Improved watcher removal deadline management with reserved tail slices for the final unsubscribe, and added retry logic for tombstone removals that fail once under transient Windows locking. * fix(history): retry failed session tree removals Tombstoned session trees whose removal fails transiently (e.g., EBUSY under Windows AV) are now re-queued in-process with bounded exponential backoff instead of sitting until the next HistoryManager construction. Prevents a single stuck tree from blocking the entire Orca process. |
||
|
|
4908a03671 |
fix: dedupe fork-copied usage history across Claude, Codex, and OpenCode scanners (#8023)
* fix: dedupe fork-copied usage history across Claude, Codex, and OpenCode scanners Coding-agent CLIs copy transcript/rollout history into new files on resume/fork, and the usage scanners deduped per-file only (or not at all), so copied history was re-counted once per descendant file (issue #8006: 38.6B tokens / $72,981 reported vs ~2.3B real). - claude-usage: cross-file turn ownership keyed on message.id:requestId; per-file ownedDedupeKeys persisted; deterministic sorted-path claim order; schema v3 -> v4 so inflated caches rebuild. - codex-usage: cross-file token_count event ownership keyed on the raw record identity (sessionId + timestamp + token tuples); fixes both the copied-prefix re-count and the total-only branch that re-counted the entire cumulative session per descendant rollout; legacy .orca-session-copies skip-bytes bridge unchanged; schema v3 -> v4. - opencode-usage: each sessionId is counted from exactly one database; the canonical opencode.db claims ahead of stale sibling copies (opencode-backup.db etc.) so backups no longer double totals, while backup-only sessions are still counted; schema v1 -> v2. Regression tests cover fork-copied files counted once (including the Codex total-only variant), duplicated OpenCode databases, and dedupe stability across cached incremental rescans. Co-authored-by: Orca <help@stably.ai> * Fix cross-tool usage double-counting for fork/resume-copied history - Widen Claude dedupe keys with message-id and uuid fallbacks so forks missing requestId still dedupe correctly, and drop Codex's sessionId from event keys since fork/resume rewrites session_meta.id while copying identical token_count records. - Track hasDeferredClaims per cached file/database across Claude, Codex, and OpenCode scanners so that when an owning file is deleted, only files that deferred a claim need reparsing to reclaim those turns instead of rescanning the entire corpus. - Let OpenCode's live opencode.db reclaim sessions from a stale backup claim once it reappears, avoiding a frozen stale snapshot. - Bump schema versions to invalidate caches built with the old, narrower ownership keys (#8006, #8013 follow-up). --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
bce3ef1776 |
Add OpenCode usage analytics (#1986)
Co-authored-by: Orca <help@stably.ai> |