mirror of
https://github.com/stablyai/orca.git
synced 2026-10-04 00:02:21 +00:00
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed, single-use, five-minute-fresh evidence. Each apply wave now samples fleet health itself right before isolation, with the monitor's evaluator, thresholds, and tolerances, for a window sized to the cell's host count (3/5/8 min), plus three lookback rules: no cell container exit in 10 min, no minute over 500 director 503s in 10 min, and director concurrency p99 within the monitor bar over 4 min. Removes the monitor-run inputs, the gate's consume/authorize steps, the break-glass override, and the same-cap-only authorization shapes in relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are unchanged. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): bound the pre-drain sample overrun and keep the drain token fresh Review follow-ups: alternating tolerated readings could hold the sample open until its step timeout, so cap the overrun at three samples past the window; record why a read failed; mint a fresh admin ID token for the drain after the sample; raise the job timeout to 90 min so a long sample cannot cancel the job past the failsafe. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule The exit rule counted every relay container exit fleet-wide, so a cell that crashes every few hours (c25, 12 a week) blocked the very roll that fixes it, and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part in. Exits are now grouped by instance, each instance is named by its own newest runtime-metrics log line, and only exits on general or migration-only cells other than the target count. An exit no configured cell can be named for trips the rule; a failed lookup is a failed read. relay-observability.tf joins the evidence-code set because the rule depends on its filter. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
31 lines
998 B
JSON
31 lines
998 B
JSON
{
|
|
"name": "@orca-cloud/relay-ops",
|
|
"private": true,
|
|
"version": "0.0.0",
|
|
"type": "module",
|
|
"main": "dist/index.js",
|
|
"scripts": {
|
|
"build": "pnpm clean && tsc -p tsconfig.build.json && node -e \"require('fs').cpSync('public','dist/public',{recursive:true})\"",
|
|
"clean": "node -e \"require('fs').rmSync('dist', { recursive: true, force: true })\"",
|
|
"dev": "tsx watch src/index.ts",
|
|
"incident:monitor": "tsx src/incident-monitor-cli.ts",
|
|
"incident:preflight": "tsx src/incident-live-preflight-cli.ts",
|
|
"incident:pre-drain-sample": "tsx src/pre-drain-sample-cli.ts",
|
|
"lint": "tsc -p tsconfig.json --noEmit",
|
|
"start": "node dist/index.js",
|
|
"test": "vitest run",
|
|
"typecheck": "tsc -p tsconfig.json --noEmit"
|
|
},
|
|
"dependencies": {
|
|
"@hono/node-server": "^1.19.17",
|
|
"hono": "^4.13.7",
|
|
"zod": "^3.25.76"
|
|
},
|
|
"devDependencies": {
|
|
"@types/node": "^24.10.0",
|
|
"tsx": "^4.21.0",
|
|
"typescript": "^5.9.3",
|
|
"vitest": "^4.1.11"
|
|
}
|
|
}
|