mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 16:02:32 +00:00
* perf(relay): batch control lease renewals per cell instead of one write transaction per host Every connected desktop renewed its own control lease with its own single-row write transaction every 30s. At ~14,000 hosts that is ~470 write transactions per second fleet-wide, each with its own transaction id, all updating the same few heap pages of relay_assignments and relay_assignment_activity_leases. Sampling three onsets at 250ms showed no lock queue and no slow statement: 60-144 backends piled into LWLock:BufferContent and Timeout/SpinDelay inside that one statement, and Query Insights attributed 152 of 157 seconds of lightweight-lock wait in the onset minute to it. The heartbeat now enqueues a due renewal instead of issuing it. A cell flushes its queue once per second, or as soon as 500 rows are waiting, through one statement that unnests the parameter arrays and applies the same CTE row-wise. Concurrent writers drop from the host count to the cell count, and transaction ids with them. Measured against a 20,000-row table: 1 row 4.1ms, 12 rows 3.3ms, 100 rows 6.6ms, 500 rows 22.7ms. Per-session semantics are unchanged. Each enqueue still resolves on a renewal and rejects with the outcome as its message, so the completed-attempt counter, the staleness guard, and every close path route exactly as before, and one outcome per row is recorded against the flush latency. Lock order is (user_id, relay_host_id), the primary key of relay_assignments, applied in JavaScript and repeated as the statement's ORDER BY. EXPLAIN confirms LockRows sits above that Sort, so a batch acquires its assignment rows in one global order. Every writer in the store locks a host's assignment row before its migration or lease rows and only ever touches one host, so a batch can only wait on a row a single-host writer holds, never the reverse. One statement also means one contended row could fail the whole batch, so a failed batch degrades to the per-host statements it replaced rather than costing every other host on the cell its renewal. * perf(relay): batch control lease renewals per cell instead of one write transaction per host Every connected desktop renewed its own control lease with its own single-row write transaction every 30s. At ~14,000 hosts that is ~470 write transactions per second fleet-wide, each with its own transaction id, all updating the same few heap pages of relay_assignments and relay_assignment_activity_leases. Sampling three onsets at 250ms showed no lock queue and no slow statement: 60-144 backends piled into LWLock:BufferContent and Timeout/SpinDelay inside that one statement, and Query Insights attributed 152 of 157 seconds of lightweight-lock wait in the onset minute to it. The heartbeat now enqueues a due renewal instead of issuing it. A cell flushes its queue every second, or as soon as 200 rows are waiting, through one statement that unnests the parameter arrays and applies the same CTE row-wise. Concurrent writers drop from the host count to the cell count, and transaction ids with them. Per-host buffer traffic is unchanged: 30 hits for one row, 25.5 per host at 12 rows, 30.1 per host at 200, against the 28.7 the single-row statement reports in production. Per-session semantics are unchanged. Each enqueue still resolves on a renewal and rejects with the outcome as its message, so the completed-attempt counter, the staleness guard, and every close path route as before, and one outcome per row is recorded against the flush latency. Row locks live until the statement commits, so a batch that waited on a contended row would hold every other row's lock for that whole wait. The assignment pass therefore takes its locks with SKIP LOCKED and reports a contended host as assignment_lock_unavailable, which the registry retries on the next tick instead of closing the control. That keeps the hold to the statement's own execution: 11.5ms for 200 rows against a 20,000-row table, and 9.4ms with a host wedged in a per-host transaction, where a blocking FOR UPDATE spends the pool's whole 1s lock_timeout and then fails every row in the flush. An unlocked present_assignment probe separates a host with no assignment row from one the skip passed over, so a skipped row can never be mistaken for a missing assignment and close a live desktop. With no wait on the assignment pass the lock order is only needed for the two later passes, and it holds: every writer takes a host's assignment row before that host's lease rows, and a host whose assignment row is held was skipped, so the batch never reaches its lease. markMigrationTargetRegistered is the one writer that locks a migration row first, and it takes no further locks. * fix(relay): answer every row of a control-renewal batch from its own lease update Review findings on the batched renewal. Two control leases on one host in one batch made the second report control_activity_not_found although both were renewed: the assignment UPDATE is offered the same target row twice, applies one source row and returns one, so the other row_index never came back. The verdict now reads renewed_lease, which has a row per input row, and the assignment UPDATE groups per host so it also stops taking an arbitrary one of the two expiries instead of the later one. The queue now partitions per (userId, relayHostId) rather than per activity, so a second control activity for one host opens the next flush instead of sharing this one. Belt to the statement fix, not a substitute: the store API has to be right for the rows it accepts. renewControlActivities recorded no outcome for a one-row flush that threw, and none at all when every row failed validation, where a mixed batch recorded its invalid_* rows. Both now record in a finally, the way the single-row path's finally always did, and the error-to-outcome mapping both paths share is one function. The four flush fields the runtime metrics event emits had no log-based metric, so add them next to the existing controlRenewalLatencyMs* entries. Applying the Terraform is a separate manual step. controlRenewalLatencyMsP50/P95/Max now measure a batched row's flush duration rather than its own statement latency. Left named as they are for history, with a line at the emit site recording that the meaning changed here.