mirror of
https://github.com/stablyai/orca.git
synced 2026-09-28 00:02:41 +00:00
* feat(cloud): add the mobile push gateway and its contract package (#8129) A small open-source service that holds the APNs key and FCM credentials and sends background push to paired phones on the desktop's behalf. Hosts authenticate with a box challenge and HMAC proof on their pairing key, the same shape the relay uses, so signed-in and accountless desktops share one path. Tokens are stored; alert text is held only for the coalescing window. The contract doc in docs/reference is the source of truth for every wire shape. The interop test runs the real desktop answerer against a real gateway-issued challenge so transcript drift fails in CI. * feat(push): register phones and send background push from the desktop (#8129) Adds the notifications.remote-push.v1 capability, the registerPush and unregisterPush RPCs on the mobile allowlist, a gateway client with a cached session and 401 re-auth, a durable unregister outbox, and a dispatcher that offers every mobile notification to the gateway after the socket fan-out. The dispatcher is fire-and-forget with one retry and drops registrations the gateway reports dead. Puts agentState on the mobile frame and fixes the #4375 wording so a working agent is never announced as finished. The relay host-proof code moves onto a shared envelope module with no behaviour change. * feat(mobile): background push registration, receive, and settings (#8129) Fetches the native APNs or FCM token, registers it with every paired host that advertises the capability, and re-registers on token change. Foreground pushes are suppressed inside handleNotification against the same seen set the socket path uses, so nothing shows twice. Taps route by host fingerprint. One Background notifications switch, off by default, with the disclaimer and needs-input / finished sub-switches; hidden until a paired desktop is new enough. Adds google-services.json and the expo-notifications plugin. * chore(cloud): Terraform and deploy workflow for the push gateway (#8129) Declares the Cloud Run service, runtime account, secrets, and orca_push database behind push_gateway_enabled, true only in production. The deploy workflow is gated like the relay's, deploys with no traffic, probes /ready and a validate-only FCM send, then shifts traffic. It runs as the shared production deploy account because the Cloud SQL rollout lease grant is foundation-owned; its extra authority is three bindings on the push service. docs/push-gateway.md carries the import commands for the resources created by hand and the APNs key rotation procedure. * docs: describe background notifications on the phone (#8129) * docs: check in the mobile push contract (#8129) Seven committed files cite it as the source of truth for every wire shape; docs/reference is allowlisted per file, so add the entry. * test(push): replay one checked-in host-proof vector on both sides (#8129) Cloud Verify installs only the cloud workspace, so the gateway suite cannot import the desktop answerer. Replace the cross-workspace import with a fixed challenge vector generated from the contract package; the gateway fixture and the desktop answerer each replay it and must produce the same HMAC. A transcript drift on either side now fails in that side's own suite. * fix(cloud): open the push gateway with invoker_iam_disabled, not an allUsers binding (#8129) The production domain-restricted-sharing policy rejects an allUsers run.invoker member, which the runbook anticipated. Opt the service out of invoker IAM the way the relay director already does; the host proof is the authentication either way. * docs(cloud): the push.onorca.dev record exists and is hand-managed (#8129) * fix(push): close review findings in the gateway (#8129) - Quota reservation takes a per-host advisory lock; READ COMMITTED admitted a whole burst past the cap (80/80 without, 60/80 with, against Postgres 16). - Challenge issuance no longer writes push_hosts; the row lands on proof verification. Stale hosts prune after 30 days. Per-IP token bucket on the two unauthenticated routes. - Streaming body limit via hono bodyLimit; a chunked body bypassed the Content-Length check. - registrationIds deduped in the schema; per-host device cap of 64; list bounded to its schema. - Gateway-side challenge TTL is the specified 10 s, not 40 s. - APNs stream settles on close as well as end/error. * fix(push): close review findings in the desktop client (#8129) - A gateway registration the registry cannot persist is enqueued for delete instead of leaking a live token. - Unregister outbox re-reads pending per pass, honours enqueues during a drain, and retries with backoff instead of waiting for the next launch. - Dispatcher batches registrations by 20 rather than starving the rest. - 401 compare-and-clear; a 401 after re-auth is unreachable; refused handshakes and 429s are cached briefly instead of re-handshaking per event. - Service is stopped on quit. * fix(mobile): close review findings in push registration and receive (#8129) - Consent generation guards a register that finishes after the switch went off; the host is re-queued for unregister instead of recorded live. - Foreground pushes seed the watermark before adopting the epoch, so a push on a never-connected session cannot wipe a valid watermark. - aps-environment follows the build via app.config.js; the iOS release workflow sets it to production. A bare plugin entry wrote development. - Pushes the OS showed while closed are marked seen before catch-up replay. - Token null result is not cached; failed capability probes are retried and never block an unregister; coalesced summaries are shown but not marked. - Unresolvable fingerprint routes nowhere and is suppressed in foreground. - Android channel ensured at boot; capability hook diffs clients by identity. * fix(cloud): harden the push deploy workflow and size the gateway to the budget (#8129) - Roll traffic back on a failed post-shift check; delete a candidate that never took traffic; retry the origin probe and the FCM probe. - Assert Terraform-owned scaling instead of mutating it from the workflow. - Build before taking the Cloud SQL rollout lease. - Declare the database pool in Terraform (2 per instance, max 2 instances) and add the gateway to the connection budget; the previous default put the shared instance 65 connections over its ceiling. - State plainly that the shared deploy identity's relay authority is inherited. * fix(push): read the runtime from shared state at push startup (#8129) Threading the runtime through launchDesktopMode put the launch module one line over the 300-line lint budget after the rebase. * fix(push): key the unauthenticated rate limit on the hop Cloud Run wrote (#8129) Cloud Run appends the connecting peer to x-forwarded-for; the limiter read the left-most value, which the caller controls, so a forged first hop earned a fresh bucket per request. * fix(push): close the final security review findings in the gateway and infra (#8129) - app.onError logs only the error name and answers a bare 500; hono's default handler printed the whole error, and a pg error carries the row in detail - a second per-IP bucket (240/min) runs ahead of the bearer lookup on every authenticated route, so forged bearers cannot spend the two-connection pool - one live session per host: minting deletes the host's earlier row - device-less hosts are pruned after 1 h, not 30 d; any keypair mints one free - notificationId is printable ASCII, since it becomes the APNs collapse header - the impersonated FCM probe token is masked in the workflow log - prevent_destroy on the Apple secrets and the orca_push database * fix(push): close the final security review findings in the desktop client (#8129) - fetch never follows a redirect: a 307 would replay the host proof and the phone's token to whatever origin the redirect named - registerPush params are strict and the paired identity is spread last - a per-device bucket (10/min) bounds a phone looping registerPush, which costs a gateway write and a synchronous registry write each time * fix(mobile): close the final security review findings in push receive (#8129) - a push with no epoch can no longer claim a seq-derived dedup key, in the foreground or from the tray; a forged seq:N could otherwise swallow the real bell at that seq - a provider-delivered push with no host catalog, or no fingerprint at all, stays unrouted instead of falling back to the hostId its raw data carries * docs(push): record the ip buckets, session and host retention, and the token-ownership limit (#8129) * fix(push): apply the schema on an untimed pool and retry statement-timeout aborts (#8129) Ports the relay's #18722 pattern to the gateway: DDL runs on a one-connection pool with statement_timeout 0 that is closed before the serving pool opens, and SQLSTATE 57014 joins the bounded transaction retry path. * fix: harden mobile push delivery and deployment recovery * feat: align mobile notification preferences with desktop delivery * fix: accept variable-length APNs device tokens * fix: deduplicate native APNs and background socket notifications
341 lines
16 KiB
YAML
341 lines
16 KiB
YAML
name: Deploy Push Gateway Production
|
|
|
|
on:
|
|
workflow_dispatch:
|
|
inputs:
|
|
confirmation:
|
|
description: Enter DEPLOY_PUSH_GATEWAY to shift production traffic
|
|
required: true
|
|
type: string
|
|
|
|
permissions:
|
|
contents: read
|
|
id-token: write
|
|
|
|
# The gateway applies its own schema at startup against the shared Cloud SQL instance, so a
|
|
# deploy is a connection-budget rollout and belongs in the same serialized group as the relay.
|
|
concurrency:
|
|
group: production-cloud-sql-rollout
|
|
cancel-in-progress: false
|
|
|
|
defaults:
|
|
run:
|
|
working-directory: cloud
|
|
|
|
jobs:
|
|
deploy:
|
|
if: >-
|
|
${{ vars.ORCA_CLOUD_OPERATIONS_ENABLED == 'true' &&
|
|
github.ref == 'refs/heads/main' }}
|
|
runs-on: blacksmith-2vcpu-ubuntu-2204
|
|
environment: production
|
|
env:
|
|
GCP_PROJECT_ID: onorca-cloud
|
|
GCP_REGION: ${{ vars.PRODUCTION_GCP_REGION }}
|
|
SERVICE_NAME: orca-cloud-push
|
|
REPOSITORY_ID: orca-cloud
|
|
IMAGE_NAME: push
|
|
PUSH_ORIGIN: https://push.onorca.dev
|
|
PUSH_RUNTIME_SERVICE_ACCOUNT: orca-cloud-push@onorca-cloud.iam.gserviceaccount.com
|
|
# Scaling the serving revision must already hold, matching push_min_instances and
|
|
# push_max_instances. Terraform owns both, and the candidate inherits them from the
|
|
# service, so this deploy never passes a scaling flag: doing so would write a
|
|
# Terraform-owned field that `lifecycle.ignore_changes` does not cover, and a later
|
|
# `push_max_instances` raise would then be reverted by every deploy. These two values
|
|
# are the expected shape, asserted before the candidate is created and again on the
|
|
# candidate itself, so a deploy that would change the gateway's Cloud SQL draw fails.
|
|
PUSH_MIN_INSTANCES: 1
|
|
PUSH_MAX_INSTANCES: 2
|
|
CONFIRMATION: ${{ inputs.confirmation }}
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
|
|
- name: Require the explicit deploy confirmation
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
test "${CONFIRMATION}" = DEPLOY_PUSH_GATEWAY
|
|
|
|
- uses: google-github-actions/auth@v2
|
|
with:
|
|
workload_identity_provider: ${{ vars.PRODUCTION_GCP_RELAY_DEPLOY_WORKLOAD_IDENTITY_PROVIDER }}
|
|
service_account: ${{ vars.PRODUCTION_GCP_RELAY_DEPLOY_SERVICE_ACCOUNT }}
|
|
|
|
- uses: google-github-actions/setup-gcloud@v2
|
|
|
|
- uses: docker/setup-buildx-action@v3
|
|
|
|
- name: Configure Docker auth
|
|
run: gcloud auth configure-docker "${GCP_REGION}-docker.pkg.dev" --quiet
|
|
|
|
# Why: the build runs before the lease. Artifact Registry is not the Cloud SQL instance,
|
|
# and a multi-minute image build inside the lease blocks every relay deploy and rehome for
|
|
# its duration. The lease below covers exactly the connection-budget window: deploy, probe,
|
|
# shift.
|
|
- name: Build and publish the immutable gateway image
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
image_tag="${GCP_REGION}-docker.pkg.dev/${GCP_PROJECT_ID}/${REPOSITORY_ID}/${IMAGE_NAME}:sha-${GITHUB_SHA}"
|
|
docker build -f apps/push/Dockerfile -t "${image_tag}" .
|
|
docker push "${image_tag}"
|
|
digest="$(gcloud artifacts docker images describe "${image_tag}" \
|
|
--format='value(image_summary.digest)')"
|
|
[[ "${digest}" =~ ^sha256:[a-f0-9]{64}$ ]]
|
|
echo "IMAGE=${GCP_REGION}-docker.pkg.dev/${GCP_PROJECT_ID}/${REPOSITORY_ID}/${IMAGE_NAME}@${digest}" \
|
|
>> "${GITHUB_ENV}"
|
|
echo "IMAGE_DIGEST=${digest}" >> "${GITHUB_ENV}"
|
|
|
|
# Held across the deploy, not just a separate schema step: the gateway opens its pool and
|
|
# applies its schema while the new revision starts, so the revision is the schema step.
|
|
- uses: ./.github/actions/cloud-sql-rollout-lease
|
|
with:
|
|
bucket: onorca-cloud-terraform-state
|
|
object: terraform/state/cloud-sql-rollout/production.lock
|
|
|
|
# Why: the candidate inherits the serving revision's scaling. A serving revision that has
|
|
# drifted below the floor would hand the candidate a cold start on every notification, and
|
|
# one that has drifted above the ceiling would hand it a larger Cloud SQL draw than the
|
|
# rollout lease was taken for. Refuse to inherit either rather than latch it.
|
|
- name: Record the serving revision and require its Terraform-owned scaling
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
serving="$(gcloud run services describe "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" --region "${GCP_REGION}" --format=json \
|
|
| jq -r '[.status.traffic[] | select((.percent // 0) > 0)]
|
|
| if length == 1 and .[0].percent == 100 then .[0].revisionName else empty end')"
|
|
test -n "${serving}"
|
|
floor="$(gcloud run revisions describe "${serving}" \
|
|
--project "${GCP_PROJECT_ID}" --region "${GCP_REGION}" \
|
|
--format="value(metadata.annotations['autoscaling.knative.dev/minScale'])")"
|
|
if [[ "${floor:-0}" -lt "${PUSH_MIN_INSTANCES}" ]]; then
|
|
echo "serving revision ${serving} holds ${floor:-0} minimum instances," \
|
|
"below ${PUSH_MIN_INSTANCES}; deploying would inherit and latch it." >&2
|
|
echo "Restore the floor first: gcloud run services update ${SERVICE_NAME}" \
|
|
"--region ${GCP_REGION} --min-instances=${PUSH_MIN_INSTANCES}" >&2
|
|
exit 1
|
|
fi
|
|
ceiling="$(gcloud run revisions describe "${serving}" \
|
|
--project "${GCP_PROJECT_ID}" --region "${GCP_REGION}" \
|
|
--format="value(metadata.annotations['autoscaling.knative.dev/maxScale'])")"
|
|
test "${ceiling}" = "${PUSH_MAX_INSTANCES}"
|
|
echo "serving revision ${serving} holds ${floor} minimum and ${ceiling} maximum instances"
|
|
echo "ROLLBACK_REVISION=${serving}" >> "${GITHUB_ENV}"
|
|
|
|
# No traffic and a per-revision tag: the candidate boots, applies schema, and is probed on
|
|
# its own URL while every phone and desktop still reaches the previous revision.
|
|
- name: Deploy the candidate revision with no traffic
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
tag="c${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}"
|
|
echo "CANDIDATE_TAG=${tag}" >> "${GITHUB_ENV}"
|
|
echo "CANDIDATE_REVISION=${SERVICE_NAME}-${tag}" >> "${GITHUB_ENV}"
|
|
gcloud run deploy "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" \
|
|
--region "${GCP_REGION}" \
|
|
--image "${IMAGE}" \
|
|
--tag "${tag}" \
|
|
--revision-suffix "${tag}" \
|
|
--no-traffic \
|
|
--quiet
|
|
candidate="$(gcloud run services describe "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" --region "${GCP_REGION}" --format=json \
|
|
| jq -er --arg tag "${tag}" \
|
|
'[.status.traffic[] | select(.tag == $tag)]
|
|
| if length == 1 then .[0] else error("tagged candidate is not unique") end')"
|
|
test "$(jq -r '.revisionName' <<< "${candidate}")" = "${SERVICE_NAME}-${tag}"
|
|
echo "CANDIDATE_URL=$(jq -r '.url' <<< "${candidate}")" >> "${GITHUB_ENV}"
|
|
|
|
# A tagged revision is directly addressable and sits outside the service-wide cap, so the
|
|
# candidate and the serving revision each draw up to the ceiling during the probe window.
|
|
# The lease is taken for exactly that doubling; a candidate that inherited a wider ceiling
|
|
# would exceed it, so the inherited scaling is asserted here too.
|
|
- name: Require the candidate to serve the exact image and inherited scaling
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
served="$(gcloud run revisions describe "${CANDIDATE_REVISION}" \
|
|
--project "${GCP_PROJECT_ID}" --region "${GCP_REGION}" \
|
|
--format='value(spec.containers[0].image)')"
|
|
test "${served}" = "${IMAGE}"
|
|
test "${CANDIDATE_REVISION}" != "${ROLLBACK_REVISION}"
|
|
candidate_ceiling="$(gcloud run revisions describe "${CANDIDATE_REVISION}" \
|
|
--project "${GCP_PROJECT_ID}" --region "${GCP_REGION}" \
|
|
--format="value(metadata.annotations['autoscaling.knative.dev/maxScale'])")"
|
|
test "${candidate_ceiling}" = "${PUSH_MAX_INSTANCES}"
|
|
|
|
- name: Probe the candidate readiness endpoint
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
[[ "${CANDIDATE_URL}" =~ ^https://[^/]+$ ]]
|
|
for attempt in $(seq 1 30); do
|
|
code="$(curl -sS -o "${RUNNER_TEMP}/push-ready.json" -w '%{http_code}' \
|
|
--max-time 10 "${CANDIDATE_URL}/ready" || true)"
|
|
if test "${code}" = 200; then
|
|
jq -e . < "${RUNNER_TEMP}/push-ready.json" > /dev/null
|
|
echo "candidate ${CANDIDATE_REVISION} is ready after ${attempt} attempt(s)"
|
|
exit 0
|
|
fi
|
|
echo "attempt ${attempt}: /ready returned ${code}"
|
|
sleep 5
|
|
done
|
|
echo "candidate ${CANDIDATE_REVISION} never reported ready" >&2
|
|
exit 1
|
|
|
|
# Why: a gateway that boots and answers /ready can still be unable to send. This proves the
|
|
# runtime account's FCM grant end to end without delivering anything: validate_only stops
|
|
# Google before any push, and the deliberately invalid token means a healthy credential
|
|
# answers INVALID_ARGUMENT. PERMISSION_DENIED is the failure this step exists to catch.
|
|
#
|
|
# Only the four verdicts below are conclusive. A 429, a 5xx, or a transport failure says
|
|
# nothing about the credential, so it is retried rather than treated as either answer; a
|
|
# denied credential still fails on the first attempt, without burning the retries.
|
|
- name: Prove the runtime identity can reach FCM
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(gcloud auth print-access-token \
|
|
--impersonate-service-account "${PUSH_RUNTIME_SERVICE_ACCOUNT}")"
|
|
test -n "${token}"
|
|
echo "::add-mask::${token}"
|
|
body='{"validate_only":true,"message":{"token":"orca-push-deploy-probe-invalid-token","notification":{"title":"Orca","body":"deploy probe"}}}'
|
|
for attempt in $(seq 1 5); do
|
|
code="$(curl -sS -o "${RUNNER_TEMP}/push-fcm.json" -w '%{http_code}' --max-time 20 \
|
|
-X POST "https://fcm.googleapis.com/v1/projects/${GCP_PROJECT_ID}/messages:send" \
|
|
-H "Authorization: Bearer ${token}" \
|
|
-H 'Content-Type: application/json' \
|
|
--data "${body}" || true)"
|
|
status="$(jq -r '.error.status // empty' < "${RUNNER_TEMP}/push-fcm.json" || true)"
|
|
echo "attempt ${attempt}: FCM validate-only send returned HTTP ${code} status ${status:-OK}"
|
|
if test "${status}" = PERMISSION_DENIED || test "${status}" = INVALID_ARGUMENT ||
|
|
test "${code}" = 401 || test "${code}" = 403; then
|
|
break
|
|
fi
|
|
sleep 5
|
|
done
|
|
if test "${status}" = PERMISSION_DENIED || test "${code}" = 401 || test "${code}" = 403; then
|
|
echo "the push runtime identity cannot send through FCM" >&2
|
|
exit 1
|
|
fi
|
|
test "${status}" = INVALID_ARGUMENT
|
|
|
|
- name: Shift all traffic to the verified candidate
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
echo "TRAFFIC_SHIFT_ATTEMPTED=true" >> "${GITHUB_ENV}"
|
|
gcloud run services update-traffic "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" \
|
|
--region "${GCP_REGION}" \
|
|
--to-revisions "${CANDIDATE_REVISION}=100" \
|
|
--quiet
|
|
serving="$(gcloud run services describe "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" --region "${GCP_REGION}" --format=json \
|
|
| jq -r '[.status.traffic[] | select((.percent // 0) > 0)]
|
|
| if length == 1 and .[0].percent == 100 then .[0].revisionName else empty end')"
|
|
test "${serving}" = "${CANDIDATE_REVISION}"
|
|
echo "TRAFFIC_SHIFTED=true" >> "${GITHUB_ENV}"
|
|
|
|
# Why: the summary is written before the origin check, not after it. Once traffic has
|
|
# moved, the rollback target is the single thing an operator needs, and a summary that only
|
|
# appeared on success would be missing in exactly the run that needs it.
|
|
- name: Publish the rollout summary
|
|
if: ${{ always() && env.CANDIDATE_REVISION != '' && env.ROLLBACK_REVISION != '' }}
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
{
|
|
echo '### Push gateway rollout'
|
|
echo
|
|
echo "Revision: \`${CANDIDATE_REVISION}\`"
|
|
echo
|
|
echo "Image: \`${IMAGE_DIGEST}\`"
|
|
echo
|
|
echo "Rollback: \`gcloud run services update-traffic ${SERVICE_NAME}" \
|
|
"--region ${GCP_REGION} --to-revisions ${ROLLBACK_REVISION}=100\`"
|
|
} >> "${GITHUB_STEP_SUMMARY}"
|
|
|
|
- name: Verify the public origin after the shift
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
for attempt in $(seq 1 30); do
|
|
code="$(curl -sS -o /dev/null -w '%{http_code}' --max-time 10 \
|
|
"${PUSH_ORIGIN}/ready" || true)"
|
|
if test "${code}" = 200; then
|
|
echo "${PUSH_ORIGIN} is ready after ${attempt} attempt(s)"
|
|
exit 0
|
|
fi
|
|
echo "attempt ${attempt}: ${PUSH_ORIGIN}/ready returned ${code}"
|
|
sleep 5
|
|
done
|
|
echo "${PUSH_ORIGIN} never reported ready after the shift" >&2
|
|
exit 1
|
|
|
|
# Why: everything after the shift runs with production on the candidate. A failure there
|
|
# is not a failure to deploy, it is a live gateway that has to go back, so the traffic move
|
|
# is undone here rather than left to whoever reads the run.
|
|
- name: Roll traffic back to the previous revision
|
|
if: ${{ (failure() || cancelled()) && env.TRAFFIC_SHIFT_ATTEMPTED == 'true' }}
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
test -n "${ROLLBACK_REVISION:-}"
|
|
gcloud run services update-traffic "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" \
|
|
--region "${GCP_REGION}" \
|
|
--to-revisions "${ROLLBACK_REVISION}=100" \
|
|
--quiet
|
|
serving="$(gcloud run services describe "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" --region "${GCP_REGION}" --format=json \
|
|
| jq -r '[.status.traffic[] | select((.percent // 0) > 0)]
|
|
| if length == 1 and .[0].percent == 100 then .[0].revisionName else empty end')"
|
|
test "${serving}" = "${ROLLBACK_REVISION}"
|
|
echo "TRAFFIC_ROLLED_BACK=true" >> "${GITHUB_ENV}"
|
|
{
|
|
echo
|
|
echo '### Push gateway rolled back'
|
|
echo
|
|
echo "Traffic returned to \`${ROLLBACK_REVISION}\`; the candidate" \
|
|
"\`${CANDIDATE_REVISION}\` no longer serves."
|
|
} >> "${GITHUB_STEP_SUMMARY}"
|
|
|
|
# Why: a candidate that never took traffic is a revision holding a warm floor and a Cloud
|
|
# SQL pool for nothing. Its tag comes off first, because Cloud Run refuses to delete a
|
|
# revision a traffic target still names, and clearing CANDIDATE_TAG makes the always() tag
|
|
# step below a no-op rather than a second failure.
|
|
- name: Delete the rejected candidate revision
|
|
if: ${{ (failure() || cancelled()) && (env.TRAFFIC_SHIFT_ATTEMPTED != 'true' || env.TRAFFIC_ROLLED_BACK == 'true') }}
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
test -n "${CANDIDATE_REVISION:-}" || exit 0
|
|
if test -n "${CANDIDATE_TAG:-}"; then
|
|
gcloud run services update-traffic "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" \
|
|
--region "${GCP_REGION}" \
|
|
--remove-tags "${CANDIDATE_TAG}" \
|
|
--quiet
|
|
echo "CANDIDATE_TAG=" >> "${GITHUB_ENV}"
|
|
fi
|
|
gcloud run revisions delete "${CANDIDATE_REVISION}" \
|
|
--project "${GCP_PROJECT_ID}" \
|
|
--region "${GCP_REGION}" \
|
|
--quiet
|
|
echo "deleted the candidate revision ${CANDIDATE_REVISION}"
|
|
|
|
- name: Drop the candidate traffic tag
|
|
if: always()
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
test -n "${CANDIDATE_TAG:-}" || exit 0
|
|
gcloud run services update-traffic "${SERVICE_NAME}" \
|
|
--project "${GCP_PROJECT_ID}" \
|
|
--region "${GCP_REGION}" \
|
|
--remove-tags "${CANDIDATE_TAG}" \
|
|
--quiet
|