Files
windmill/docs/git-sync-pull-design.md
T

18 KiB

Design: Automatic git → Windmill sync (pull-based)

Status: draft for review — exploration on branch explore-git-sync-improvements

1. Problem

Git sync is guided and automatic in one direction only. Windmill → repo is fully managed: every deploy enqueues a deploymentcallback job that runs the hub sync script and commits to the repo. The reverse direction (repo → Windmill) requires each customer to install a GitHub Action that runs wmill sync push with a long-lived Windmill token stored as a repo secret, against an instance URL that GitHub-hosted runners must be able to reach.

This breaks down because each customer runs their own instance:

Customer setup GH runner → instance GitHub webhook → instance Instance → GitHub
Windmill cloud
Self-hosted, public URL
Self-hosted behind VPN/firewall
GHES on the same private network (github.com runners)

Only instance → GitHub outbound works for everyone. The current GH Action design sits in the worst column; it also requires manual workflow-file installation, token provisioning, and ongoing maintenance per customer.

2. Current state (what already exists)

The building blocks are mostly shipped:

  • Instance-side pull already works. PullWorkspaceModal.svelte runs the hub init script (hubPaths.gitInitRepo) with pull: true, dry_run preview, and use_promotion_overrides as a worker job. It clones the repo with the repo credentials and applies the diff to the workspace. It is manual-only today.
  • Managed GitHub App. windmill-sync-helper (github.com / *.ghe.com) is installed by the customer; installation tokens are minted by the customer portal (windmill-customer-service, route /github_sync/token) against a JWT the instance stores per installation (workspace_settings.git_app_installations, token logic in backend/windmill-common/src/git_sync_ee.rs). The portal is stateless: it keeps no installation registry and receives no webhooks.
  • Self-managed app for GHES. Customers create their own GitHub App; app id + private key live in instance settings (GhesAppSettings.svelte, get_self_managed_installation_token). The current setup checklist tells the customer to leave the app webhook inactive.
  • Webhook machinery. The native GitHub trigger already creates/deletes repo webhooks via the REST API (windmill-native-triggers/src/github/external.rs) and GitHub HMAC (X-Hub-Signature-256) verification exists in windmill-trigger-http/src/http_trigger_auth.rs.
  • Loop prevention convention. Windmill-authored commits carry a [WM] prefix so CI can ignore them.
  • Promotion plumbing. The CLI implements --promotion <branch> resolving promotionOverrides from wmill.yaml; the pull job already accepts use_promotion_overrides. PR creation on wm_deploy/** branches is done by a documented GH Action (gh pr create), not by Windmill.

What is missing is only the trigger (push event → pull job), the routing (event ref → workspace), and PR creation moving instance-side for promotion/fork parity.

3. Goals / non-goals

Goals:

  • Zero-CI automatic repo → Windmill deployment: install app, pick repo, done.
  • Works for every connectivity profile, degrading gracefully from instant (webhook) to near-real-time (polling).
  • Cover all documented setups: basic sync, multi-repo, promotion mode (single- and cross-instance), workspace forks, local-dev git entry.
  • No new credentials handed to GitHub (no Windmill tokens as repo secrets).
  • Coexist with customer CI: customers who keep the GH Action lose nothing.

Non-goals (this design):

  • Replacing customer CI pipelines (tests, lint, custom gates).

  • A central event-relay through the portal. Evaluated and rejected for v1: the portal stays a stateless token minter. (Revisit only if private instances need sub-minute latency; see §12.)

    Rationale — the proxy's only real advantages are (a) zero-friction enablement for existing installs (app-level push events need no new permissions, unlike repo-webhook creation) and (b) central delivery observability. Against that: it does not improve reachability (the portal forwards to the same instance URL GitHub would deliver to, so private instances need polling either way); app-level events fire for every push to every installed repo — org-wide — so portal cost scales with customer push volume, not synced repos; the portal becomes stateful and availability-coupled; full push payloads (commit messages, author emails) transit Windmill infrastructure; and the instance cannot verify portal signatures, whereas repo webhooks get per-repo HMAC. The permission-bump friction is mitigated contextually (approval link shown in settings when enabling auto-pull) and polling covers the gap until approved.

  • GitLab/Bitbucket/Azure DevOps parity. The polling tier covers them credential-wise; their webhook tiers are follow-ups.

4. Design overview

flowchart LR
  subgraph GitHub
    R[Repo] -- "push event" --> W[repo webhook<br/>created via API]
  end
  subgraph Instance
    W -- "HMAC-verified POST" --> E["/api/w/:ws/github_app/push_webhook"]
    P[poller<br/>git ls-remote / schedule] --> REC
    E --> REC[reconcile:<br/>ref match? sha moved?<br/>debounce]
    REC --> J[pull job<br/>existing hub script, pull:true]
    J --> WS[(workspace)]
  end

One principle drives the security model: webhooks and polls are hints, the pull is authoritative. A trigger never carries content; it only causes the instance to compare the remote HEAD against last_synced_sha using its own credentials and enqueue the existing pull job if the branch moved. A forged or replayed trigger can only cause a cheap no-op reconcile.

5. Trigger tier 1 — webhooks

5.1 Managed app (github.com, GHE Cloud orgs)

A GitHub App's own webhook URL is fixed app-wide (it points at the portal), so per-instance delivery uses repository webhooks created dynamically with the installation token (POST /repos/{owner}/{repo}/hooks). Requires adding the Repository webhooks: read & write permission to windmill-sync-helper (see §9).

Flow when a repo is connected (or auto-pull is enabled on an existing repo):

  1. Instance generates a per-repo secret, stored in the repo's git-sync settings.
  2. Instance creates the webhook via API: events ["push"] (later "pull_request", §11), URL {base_url}/api/w/{workspace}/github_app/push_webhook/{repo_settings_id}, secret set.
  3. GitHub immediately delivers a ping event. If the ping is not received within ~10 s, the instance deletes the hook and falls back to polling, surfacing "instance not reachable from GitHub — using polling (interval Xm)" in the UI. This doubles as an automatic reachability test; no guessing about firewalls.
  4. Incoming deliveries are verified with X-Hub-Signature-256 against the stored secret (reuse the HMAC code from http_trigger_auth.rs), then handed to the reconciler (§7).

Webhook lifecycle: deleted when the repo is disconnected or auto-pull disabled (reuse the delete pattern from workspace_integrations.rs); recreated on settings change; orphan hooks are detectable via GET /repos/.../hooks filtered by our URL prefix.

5.2 Self-managed app (GHES, *.ghe.com data residency)

The customer owns the app, so the app-level webhook can point directly at the instance — no repo hooks needed, one webhook covers all installed repos:

  • GhesAppSettings gains a "Webhook secret" field; the setup checklist changes from "Uncheck Active under Webhook" to "set webhook URL to {base_url}/api/github_app/webhook (instance-global endpoint), subscribe to Push events, paste this generated secret".
  • The instance-global endpoint verifies HMAC with the instance-level secret and routes by repo full name + installation to matching workspaces (the routing data is in workspace_settings).
  • GHES typically shares a network with the instance, so this works air-gapped — the hardest github.com case is the easiest GHES case.

Polish (optional, recommended): replace the 8-step manual app-creation checklist with the GitHub App Manifest flow (supported on GHES), which pre-configures permissions, events, and webhook URL/secret in one click and eliminates checklist drift.

6. Trigger tier 2 — polling fallback

For private instances, plain token/SSH credentials (no app), and permission-not-yet-approved installs:

  • Per auto-pull-enabled repo, on a configurable interval (default 60 s, surfaced in settings), run git ls-remote <repo> <tracked refs> and compare against last_synced_sha per ref. ls-remote is a single cheap round-trip; no clone.
  • Implemented as an internal scheduled task keyed by repo settings id (not a user-visible schedule), running on workers like other background jobs.
  • Polling is also the safety net under webhooks (missed deliveries, GitHub outages): when a webhook is active, the poll interval relaxes (e.g. 10 min) rather than turning off. This is the ArgoCD model: poll for correctness, webhook for latency.

7. Reconcile and pull semantics

Routing — an event/poll result is (repo, ref, head_sha, sender):

  • Tracked branch (basic sync): ref equals the repo settings branch → pull into the owning workspace.
  • Promotion target branch: ref equals a repo's promotion target → pull with use_promotion_overrides: true (resolves promotionOverrides from wmill.yaml, mirroring CLI --promotion).
  • Fork branches wm-fork/<parent-branch>/<fork-name>: parse the branch name, route to the fork workspace if it exists; ignore (log) if not.
  • Anything else: ignore.
  • Matching keys on repo + workspace_id — never base_url (avoids the known internal-vs-public URL mismatch that breaks CLI workspace matching).
  • One repo may match several workspaces (team partitioning with different path filters): fan out, each workspace pulls with its own filters; wmill.yaml in the repo stays authoritative for include/exclude.

Loop prevention (pull → deploys → deployment callback → commit → push event):

  1. Skip events whose sender is the app bot (windmill-sync-helper[bot] / the GHES app's bot) or whose head commit message carries the [WM] prefix.
  2. Compare head_sha to last_synced_sha before enqueuing; the pull job records the synced sha on success.
  3. Deploys applied by the pull job are tagged so the deployment callback can skip the no-op commit (belt and braces — the diff should be empty anyway).

Concurrency and debounce:

  • At most one pull job per (repo, workspace) at a time; rapid pushes coalesce (same pattern as the 5 s deploy-callback batching, keyed on repo).
  • Pulls and in-flight Windmill → repo commits on the same branch serialize on the same key to avoid races.

Execution: reuse the existing pull path of the hub init script for v1. Two known fragilities to fix as part of making this an unattended automation:

  • The hub script pins a windmill-cli version; pin lag has caused 422s against newer backends. The automated path must pin the CLI to the instance's own version (or the pull logic moves into the backend natively as a v2).
  • Failures must be visible: pull jobs appear in the runs list like deployment callbacks today (/runs?job_kinds=deploymentcallbacks equivalent), plus a per-repo "last sync" status chip in git-sync settings and the workspace error handler firing on repeated failures.

8. Settings and schema

Per-repo (GitRepositorySettings, workspace_settings.git_sync):

auto_pull:           Option<AutoPullSettings>
  enabled:           bool
  mode:              "webhook" | "polling" | "auto"   (auto = try webhook, fall back)
  poll_interval_s:   Option<u32>                       (default 60; 600 when webhook active)
  webhook_id:        Option<i64>                       (GitHub hook id, managed app)
  webhook_secret:    Option<String>                    (encrypted; per-repo, managed app)
  last_synced_sha:   Option<map<ref, sha>>             (per tracked ref)
  last_pull_status:  Option<...>                       (sha, time, job id, error)

Instance-level (GHES app config): webhook_secret alongside app id/private key.

No new tables; everything extends existing JSONB settings. A migration is only needed if we decide last_synced_sha/status churn doesn't belong in workspace_settings (alternative: small git_sync_pull_state table keyed by workspace + repo path — decide at implementation review).

New instance endpoints (EE):

  • POST /api/w/{workspace}/github_app/push_webhook/{repo_settings_id} — per-repo hook receiver (managed app), HMAC-verified, returns 202.
  • POST /api/github_app/webhook — instance-global receiver (self-managed app), HMAC-verified, routes internally.
  • Both are unauthenticated-but-verified endpoints; rate-limited; bodies are treated as hints only (§4).

9. GitHub App permission migration

New permissions needed, bundled into one update (each update re-prompts every existing installation's org admin):

Permission Used for Phase
Repository webhooks: write dynamic repo hook create/delete 1
Pull requests: write instance-side PR creation (promotion/forks) 2
Checks: write PR diff preview checks 3

Rollout behavior: until an org approves, webhook creation fails with a distinguishable error → the instance shows "approval pending" and stays on polling. Nothing breaks; latency is the only cost. The self-managed (GHES) checklist/manifest gains the same permissions — no central approval involved.

10. Setup UX

The git-sync wizard (DetectionFlow / repo card) gains a third guided direction after the existing "test connection" and "initialize repo" steps:

  • Toggle: "Automatically deploy changes from Git" (per repo).
  • On enable: try webhook (ping self-test) → show resulting mode and latency ("instant via webhook" / "polling every 60s — instance not reachable from GitHub"), with a re-test button.
  • The success modal's current "set up GitHub Actions" doc link becomes the advanced/CI path, not the default instruction.
  • Promotion-mode repos additionally show "PRs will be opened by Windmill" once §11 lands.

11. Promotion & forks parity: PR creation moves instance-side

Today two documented GH Actions exist only to call gh pr create:

  • on push to wm_deploy/** (promotion mode),
  • on fork-branch commits (workspace forks).

Both move into the deployment callback: after pushing the branch, open (or reopen) the PR via the installation token (POST /repos/.../pulls), targeting the promotion/parent branch. Benefits: one less customer-installed workflow, and the deploy UI can link directly to the PR. The merge side is already covered by §7 routing (push event on the target branch). The documented actions remain valid for customers who want CI in the path.

12. Later: PR diff preview checks

With pull_request events (webhook tier) and checks: write: on PR opened/synchronized, run the existing dry_run: true pull and post the diff summary as a check run. This replicates the CI dry-run preview with zero customer CI and completes the "Cloudflare Pages" experience: install app → merges deploy, PRs show a Windmill diff. Private instances without webhooks can optionally get this via portal event polling — the only scenario that would justify reopening the portal-relay idea (portal stores events; instances poll /github_sync/events with their existing installation JWT). Explicitly out of scope for v1.

13. Coverage vs documented setups

Documented setup Covered by
Basic git sync (workspace ↔ branch) §5/§6 trigger + existing pull job
Multi-repo primary/secondary per-repo toggle; secondaries stay push-only
Promotion mode, single instance §7 promotion routing + §11 PR creation
Promotion mode, cross-instance each instance triggers independently — strictly better than CI (no cross-instance tokens/URLs)
Workspace forks (wm-fork/**) §7 fork routing + §11 — ship last, most edge cases
PR dry-run preview §12 (optional follow-up)
Local dev, git as entry point ordinary push events; nothing special
Customers with real CI gates unchanged; pull triggers are idempotent and coexist

14. Rollout plan

  1. Phase 1 — polling + reconcile + settings + UX (no app changes): covers every customer immediately, including air-gapped GHES and plain-token repos. Fix CLI version pinning for the unattended path.
  2. Phase 2 — webhooks: managed-app permission bump (all three permissions at once), dynamic repo hooks with ping self-test; GHES app-level webhook + secret field + checklist/manifest update.
  3. Phase 3 — PR creation instance-side (promotion/fork parity), docs updated to demote the PR-creation actions to optional.
  4. Phase 4 — PR diff checks (and only then, if demanded, portal event polling for private instances).

15. Open questions

  • Does last_synced_sha/pull-status churn stay in workspace_settings JSONB or move to a dedicated table? (Write frequency vs settings-blob contention.)
  • Exact debounce window for pull coalescing (reuse 5 s like deploy callbacks, or longer since clones are heavier?).
  • Should phase 1 polling default to on for newly connected repos, or strictly opt-in? (Opt-in proposed; revisit after adoption data.)
  • GHE Cloud orgs with IP allowlists: confirm hook deliveries to customer instances aren't filtered; the ping self-test catches it operationally either way.
  • v2: move pull execution from the hub script into the backend natively (removes CLI pinning and hub round-trip entirely)?