mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 16:02:32 +00:00
* feat(relay): count failed cell-inventory lock acquisitions The cell inventory lock is taken NOWAIT, so contention errors with 55P03 and retries instead of waiting. CellInventoryHoldSamples.record only runs after a successful acquisition, so the hold metrics were structurally blind to the dominant failure mode: production showed ~65 failed fleet-wide acquisitions per minute while cellInventoryHoldMsMax read a benign 53ms mean. Count failures next to the holds and publish them as cellInventoryLockUnavailable in orca_relay_runtime_metrics. Drained on both the commit and the rollback path, since a 55P03 rolls its transaction back. * fix(relay): separate request-path lock timeouts from sweep deferrals Review caught that the first counter only incremented under failIfUnavailable, which is the sweep mode. Background sweeps take the inventory NOWAIT and re-derive a skipped candidate next tick, so those deferrals are by design and already reported as orca_relay_sweep_cell_inventory_busy. The request path uses a bounded lock_timeout instead, whose expiry raises the same 55P03 without NOWAIT and was not counted at all -- so the metric measured only the benign population and missed the user-visible one. Split them: cellInventoryLockUnavailable for NOWAIT deferrals, cellInventoryLockTimeouts for expired bounded waits. Production over 30 minutes shows why the distinction matters -- roughly 1,200 fleet-wide sweep deferrals against roughly 10/min request-path timeouts. Adds transaction-path coverage for both drains, which were previously unpinned. Timeouts count per attempt, not per request, since 55P03 is retryable. * fix(relay): publish the cell-inventory lock metrics to Cloud Monitoring google_logging_metric.relay_snapshot only creates metrics for fields listed in relay_runtime_metrics, and the cellInventoryHold* fields were never added when the hold telemetry landed. They have been log-only since, so nothing could alert on the lock and the contention stayed invisible in exactly the way the telemetry was meant to prevent. Maps the three hold fields and both new failure counters. Also corrects the field comment: the split is by wait policy, not by caller. assignOnce takes the inventory fail-fast on its first placement attempt, so request-reachable sites land in cellInventoryLockUnavailable too; that lane reads as contention pressure, and the expired bounded wait is the stall lane.