fix(orchestration): retry worker_done while the Orca runtime is briefly unreachable (#23984)

* fix(orchestration): retry worker_done while the Orca runtime is briefly unreachable

A worker reports worker_done once and ends its turn, so a few-minute app
outage silently stranded finished work at dispatched. The CLI now retries
worker_done on runtime_unavailable for about two minutes with backoff,
reusing one request id so the host's mutation ledger replays rather than
double-applies it, then prints the existing recovery command.

The contract probe no longer caches a failed status.get, which otherwise
made every retry fail without reaching the app.

Part of STA-8833.

* refactor(cli): pass the worker_done retry window in the mutation options bag

* Revert "refactor(cli): pass the worker_done retry window in the mutation options bag"

The options bag is forwarded to client.call as-is; the retry window is not a client.call option, and folding it in needed a value scan to keep the no-options call shape.

* test(cli): fold the explicit retry-request case and drop a vacuous timing assert

* fix(cli): keep worker_done recovery when the last retry fails before sending

Also skip the Unix-socket retry test on Windows and remove its temp profile.
This commit is contained in:
Jinwoo Hong
2026-09-30 13:04:57 -04:00
committed by GitHub
parent e03870403e
commit 46d6b76ae8
6 changed files with 380 additions and 72 deletions
+7 -1
View File
@@ -245,7 +245,13 @@ export class RuntimeClient {
private async ensureOrchestrationContractCompatible(timeoutMs: number): Promise<void> {
if (!this.orchestrationContractCheck) {
this.orchestrationContractCheck = this.checkOrchestrationContractCompatibility(timeoutMs)
this.orchestrationContractCheck = this.checkOrchestrationContractCompatibility(
timeoutMs
).catch((error: unknown) => {
// Why: a failed probe must not be cached, or a retry after a brief outage never reaches the app.
this.orchestrationContractCheck = null
throw error
})
}
await this.orchestrationContractCheck
}