mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 00:02:29 +00:00
* fix(relay): commit the cell counter in one round trip; cells boot without the DB Step 2 (option B) cell image: - One-round-trip counter commit at acquireActivity, releaseActivity and activateControl: the final counter UPDATE and COMMIT go as one simple-query message. Server errors mean COMMIT never ran (retry as today; 22012 = no row, rolled back and disambiguated outside the transaction); a lost connection is never retried. - Cells skip the schema apply and region backfill, so they listen while the database is down and turn ready on their first successful query. - G13: rehome target connection headroom folded into the existing NOWAIT UPDATE, excluding the host's own reservation by key. - fixLevel on every runtime metrics line, plus declared (not applied) cell fix-level metrics and alert. - Per-desktop drain disconnect-gap measurement from existing log lines. - Census test that fails on floating database promises; fixes two shutdown sites. Lock-wait sample keeps the combined role. * fix(relay): make the outdated-image alert creatable: one PromQL condition, 1 h lookback, fixed floor A PromQL condition must be the only condition in its policy, and alerts on log-based metrics may look back at most 25 h. Replace the 6-day/7-day design with relay_cell_min_fix_level (tfvars, raised by a targeted apply after each wave) and one query: a serving cell below the floor or reporting no level, sustained 6 h. Drops the separate without-level metric. * fix(relay): review fixes: gap-script ordering, wider promise census, fused-path guard, row-busy as scheduled - Drain gap script: sort closes by time (gcloud exports newest first) and refuse an invalid drain start. - Census: any floating promise in relay src, including callback-discarded and never-read ones, with a reviewed never-rejects list. - Test the fused counter commit through the store the server builds, so a wrapper that stops forwarding commitWithFinal fails CI. - Same-cap shadow gate: a row-busy refusal (the host's own release still holds its row) is a scheduled 503, like an own early retry. No client change. * test(relay): judge drain redials by no host refused twice, not a refusal count The row-busy count tracks how many releases are still in flight at the dial (80 of 180 every run at 1 s, against a bar of 90). What matters is that the release has finished by the next dial: assert no host is refused twice, keep the time-to-placed p95 bound. * fix(relay): cap row-busy as scheduled at the drain-return admissions; bound the gap script's window Shadow gate: a row-busy refusal of a drained host follows its drain-return lane admission, so per minute only that many (plus a rounding margin of 2) are scheduled; the rest stay non-drain, so row contention the drain does not explain still fails the budget. Gap script: --drain-ended-at excludes the new container's closes after the roll; later grants still close a gap.
58 lines
2.5 KiB
TypeScript
58 lines
2.5 KiB
TypeScript
import type { RelayDatabase } from './database.js'
|
|
import type { DatabaseLockWaitSample } from './relay-observability.js'
|
|
|
|
// Every relay process connects as the same user through a socket, so neither
|
|
// pg_stat_statements nor Query Insights can tell a director's lock wait from a
|
|
// cell's. application_name (`orca-relay/<role>/<cell>`) can, so this samples it.
|
|
// The holder is the root of the wait chain: in a row-lock convoy every later
|
|
// waiter is blocked by the first waiter, not by the transaction holding the row.
|
|
// Blockers are read once per waiter; non-relay waiters stay in so chains through
|
|
// them resolve, and the depth cap bounds a cycle. The table is the first relay
|
|
// table the waiting statement names, which may be one it only references.
|
|
// It shares the director's 3-slot pool, so when every slot is a lock waiter the
|
|
// sample queues and undercounts director waiters; a dedicated connection fixes that.
|
|
const LOCK_WAIT_SAMPLE_SQL = `
|
|
WITH RECURSIVE waiting AS MATERIALIZED (
|
|
SELECT pid, application_name, query, (pg_blocking_pids(pid))[1] AS blocker
|
|
FROM pg_stat_activity
|
|
WHERE datname = current_database() AND wait_event_type = 'Lock'
|
|
), chain AS (
|
|
SELECT pid AS waiter, blocker AS pid, 1 AS depth FROM waiting
|
|
UNION ALL
|
|
SELECT chain.waiter, waiting.blocker, chain.depth + 1
|
|
FROM chain JOIN waiting ON waiting.pid = chain.pid
|
|
WHERE chain.depth < 8
|
|
), root AS (
|
|
SELECT DISTINCT ON (waiter) waiter, pid FROM chain ORDER BY waiter, depth DESC
|
|
)
|
|
SELECT split_part(w.application_name, '/', 2) AS waiter_role,
|
|
COALESCE(
|
|
substring(w.query FROM '\\m(relay_cells|relay_assignments)\\M'),
|
|
'other'
|
|
) AS waited_table,
|
|
split_part(holder.application_name, '/', 2) AS holder_role,
|
|
COUNT(*) AS waiters
|
|
FROM waiting w JOIN root ON root.waiter = w.pid
|
|
LEFT JOIN pg_stat_activity holder ON holder.pid = root.pid
|
|
WHERE w.application_name LIKE 'orca-relay/%'
|
|
GROUP BY 1, 2, 3`
|
|
|
|
const RELAY_ROLES = new Set(['director', 'cell', 'combined'])
|
|
|
|
export async function readPostgresLockWaitSample(
|
|
database: RelayDatabase
|
|
): Promise<DatabaseLockWaitSample> {
|
|
const rows = await database.query(LOCK_WAIT_SAMPLE_SQL)
|
|
return rows.map((row) => ({
|
|
waiterRole: relayRole(row['waiter_role']),
|
|
table: String(row['waited_table']),
|
|
holderRole: relayRole(row['holder_role']),
|
|
waiters: Number(row['waiters'])
|
|
}))
|
|
}
|
|
|
|
// Anything else (an operator session, a finished holder) stays one bounded key.
|
|
function relayRole(value: unknown): string {
|
|
return typeof value === 'string' && RELAY_ROLES.has(value) ? value : 'other'
|
|
}
|