Files
orca/cloud/apps/relay/src/postgres-lock-wait-sample.ts
T
Jinwoo Hong 10bea3a8fa feat(relay): report lane service times, 503 causes, per-site hold p99 and a director-vs-cell lock-wait split (#25645)
* feat(relay): report lane service times, 503 causes, per-site hold p99 and lock-wait split

Adds the step-1 observability: drain-return and sticky slot service times,
every /v1/assign 503 by cause with a non-drain total, a p99 per lock hold
site (drain-return regional rows get their own site), a director-side
pg_stat_activity sample that splits lock waiters by director vs cell, and
a log line for each reserved_requests drift reconciliation corrects.
Log-based metrics for the new fields are declared in Terraform, not applied.

* fix(relay): classify a lock waiter by the first relay table its statement names

* fix(relay): attribute lock waiters to the root holder; document the director hold alert change

Every waiter after the first in a row-lock convoy is blocked by the first
waiter, so the sample now walks pg_blocking_pids to the root, reading it once
per waiter. The sampler is single-flight. Drain-return regional holds feeding
the non-paging director hold policy is documented as expected.

* docs(relay): note the lock-wait sample undercounts director waiters when the pool is full
2026-10-05 17:55:52 -04:00

58 lines
2.5 KiB
TypeScript

import type { RelayDatabase } from './database.js'
import type { DatabaseLockWaitSample } from './relay-observability.js'
// Every relay process connects as the same user through a socket, so neither
// pg_stat_statements nor Query Insights can tell a director's lock wait from a
// cell's. application_name (`orca-relay/<role>/<cell>`) can, so this samples it.
// The holder is the root of the wait chain: in a row-lock convoy every later
// waiter is blocked by the first waiter, not by the transaction holding the row.
// Blockers are read once per waiter; non-relay waiters stay in so chains through
// them resolve, and the depth cap bounds a cycle. The table is the first relay
// table the waiting statement names, which may be one it only references.
// It shares the director's 3-slot pool, so when every slot is a lock waiter the
// sample queues and undercounts director waiters; a dedicated connection fixes that.
const LOCK_WAIT_SAMPLE_SQL = `
WITH RECURSIVE waiting AS MATERIALIZED (
SELECT pid, application_name, query, (pg_blocking_pids(pid))[1] AS blocker
FROM pg_stat_activity
WHERE datname = current_database() AND wait_event_type = 'Lock'
), chain AS (
SELECT pid AS waiter, blocker AS pid, 1 AS depth FROM waiting
UNION ALL
SELECT chain.waiter, waiting.blocker, chain.depth + 1
FROM chain JOIN waiting ON waiting.pid = chain.pid
WHERE chain.depth < 8
), root AS (
SELECT DISTINCT ON (waiter) waiter, pid FROM chain ORDER BY waiter, depth DESC
)
SELECT split_part(w.application_name, '/', 2) AS waiter_role,
COALESCE(
substring(w.query FROM '\\m(relay_cells|relay_assignments)\\M'),
'other'
) AS waited_table,
split_part(holder.application_name, '/', 2) AS holder_role,
COUNT(*) AS waiters
FROM waiting w JOIN root ON root.waiter = w.pid
LEFT JOIN pg_stat_activity holder ON holder.pid = root.pid
WHERE w.application_name LIKE 'orca-relay/%'
GROUP BY 1, 2, 3`
const RELAY_ROLES = new Set(['director', 'cell'])
export async function readPostgresLockWaitSample(
database: RelayDatabase
): Promise<DatabaseLockWaitSample> {
const rows = await database.query(LOCK_WAIT_SAMPLE_SQL)
return rows.map((row) => ({
waiterRole: relayRole(row['waiter_role']),
table: String(row['waited_table']),
holderRole: relayRole(row['holder_role']),
waiters: Number(row['waiters'])
}))
}
// Anything else (an operator session, a finished holder) stays one bounded key.
function relayRole(value: unknown): string {
return typeof value === 'string' && RELAY_ROLES.has(value) ? value : 'other'
}