mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 08:02:12 +00:00
My previous commit inferred it from "the transaction already took a fleet-wide lock", and a second review round demolished that three ways. The inference is unsound: a fleet-wide lock does not fence rows another transaction creates after it, and the interleaving deadlocks. Over an empty inventory it locks nothing at all while still reading as total cover, and that one deadlocks at bootstrap, where nothing else serialises registration. And it was self-disarming, which is the part that decided this. The predicate could not distinguish a one-row `cell_id IN (?) ORDER BY cell_id ASC` from the fleet lock -- so `lockCellRows`, the helper the per-cell conversion moves every site onto, would have set the exemption and silenced the check exactly when it started mattering. A guard that switches itself off the moment its subject arrives is worse than an acknowledged gap, because its silence still reads as evidence. So the upsert goes back to set-extension and the hazard is written down instead: in the source, and as a test that asserts the gap and will fail when someone closes it. Closing it needs create-versus-collide made explicit -- a flag threaded from the three registration sites, or `RETURNING (xmax = 0)`, which is Postgres only and would cost the both-dialects property. That belongs with the conversion, which is when the hazard becomes reachable. It is latent today because every upsert site takes the fleet-wide lock first. Also replace the single-row predicate, which produced FALSE reports on correct SQL -- `WHERE cell_id = ? AND (enabled = 1 OR enabled = 0)` read as multi-row, and those throw in tests. The unordered-lock check now asks only whether a locked read constrains cell_id at all, and no longer consults how many rows came back, so it is finally shape-based rather than fixture-dependent. Conservative on purpose: under-reporting costs a catch, over-reporting costs trust in the silence.