mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-09-08 16:03:27 +00:00
* feat: durable dbt state per environment, and `--defer` onto it `dbt retry` worked off two artifacts and only one was durable: `dbt_run_state` holds `run_results.json` keyed by principal, and the manifest lived on worker-local disk under a four-generation cache. That is enough to resume the last run and nothing else — the next run of a project usually lands on a worker holding neither artifact — so deferral had nothing to read. Adds `dbt_environment_state`: one row per (workspace, script path, environment), holding `manifest.json` and `run_results.json` from the last successful run, with the blob inline under `DBT_STATE_INLINE_MAX_BYTES` and in the workspace's object storage above it. Environment is the warehouse, the target, and the database and schema they resolve to, so a repointed warehouse or a moved schema reads as an environment nothing has published rather than as state whose relation names no longer fit. A run publishes it when its graph becomes what the script owns and it succeeded — the same condition, and the same reason: an invocation that scoped its own model set describes where the caller put those relations, not where the project's models live. `defer` is a `build` command-block field defaulting to the descriptor's own, and the state is materialised into the job directory for `--defer --state`. The retry path already did that materialisation for `dbt retry`; both go through one `write_state_dir` now. `--state` is also where `dbt retry` reads the run it resumes, so a retry on dbt-core 1.x takes `--defer-state` instead, and one on an engine without that flag is refused before the build rather than rebuilding its nodes with every unbuilt `ref()` resolving into the schema this run writes into. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ned2pmRJwB3GpenEcrA9TF * fix: address the local review of the dbt environment state The oversized-artifact home moves from the workspace's object storage to the instance's, where every other internal worker artifact already lives. The workspace bucket is the one members read and write through `job_helpers/*` with a caller-supplied key and only `volumes/` is reserved there, so a manifest under it is one any member could replace — and the next deferring run would hand dbt an attacker-chosen `defer_relation` for every unbuilt `ref()` while holding the script's warehouse credentials. The environment key takes the target dbt actually runs rather than the descriptor's `profile.target`, which is absent whenever the target is inherited from the workspace warehouse or the project's own `profiles.yml` — filing every inherited target under one empty name, while a `target.name` macro decides where a model is built. `write_profiles` returns a named struct now that it resolves one more thing. Publishing takes the row's lock before uploading, so two publishers of one environment cannot interleave their uploads and leave one run's manifest beside another's results, and carries the live-dbt-script guard the retry state already had, so a job finishing after its script was renamed, archived or deleted cannot recreate state at a path for whatever is created there next. A rename now clears the environment state instead of moving it: an oversized artifact's key is derived from the path, so a moved row would keep pointing at a key a script created at the old path publishes over. A build recovered by the automatic in-job node retry publishes its manifest without results — `run_results.json` is then the retry's, naming only the nodes it redid — and the refusal for an environment with nothing published names the runs that cannot publish rather than suggesting a run that would not help. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: serialize dbt state publishers on an advisory lock The row lock only serializes publishers once a row exists, and the first publish of an environment — two runs of a newly deployed script — is exactly when two of them are most likely to race and interleave their uploads. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: make dbt state publication atomic and bind it to the version that ran Every publication now writes its own object keys and the row switches to them in one statement, so an upload never overwrites an artifact the committed row still names: a run failing between its two uploads, or between them and its row, leaves the state pointing at the pair it already had. The objects a commit displaces are dropped afterwards — never before, since a reader that has already read the row is about to fetch them — and a reader that loses that race re-reads the row once rather than reporting a state that is there. What a publication uploaded and then could not commit is dropped on the way out. The write's guard names the VERSION rather than the path: the live dbt script there must be the one this job ran, or a later version of it. "Some live dbt script is here" is also satisfied by a script created at a path this one was renamed away from, and this job's manifest would then become that project's deferral state. A preview names no version and so publishes nothing. A `show` defers too. It compiles the model it previews, so a model whose upstream this environment built and this run did not is exactly the case a deferral exists for, and every engine takes the flags on it. Three comments said "the workspace's object storage" where the code deliberately uses the instance's, which is the whole security argument; `mib()` labelled MiB values MB. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: hold the script row across a dbt state publication, and let a rename move it The version guard read `script` without a lock, so lifecycle cleanup could find no environment row to clear, finish, and leave this transaction to commit state at a path a new script goes on to occupy. It now holds that row (`FOR SHARE`) for the rest of the publication — taken before the sidecar, the order every other dbt writer takes — and the artifacts are uploaded before the transaction, so the lock covers the row work rather than a network round trip. A commit that reports an error may still have committed: what was lost can be the acknowledgement. Dropping this run's objects then leaves the committed row naming objects that are gone, so an orphan is the cheaper side to take. A failed second upload left the manifest it had already written behind; it is dropped now. Per-publication keys retired the reason a rename cleared the environment state rather than moving it: the path is only a prefix, and the row is what names an artifact, so a script created at the old path can no longer publish over a moved row. The rename moves both halves again. `dbt ls` gets the deferral flags too, without which a `result:` selector — which reads `run_results.json` out of the state directory, and which `select` passes to dbt verbatim — fails before the build that would have honoured it. Also: the migration was the last site describing the workspace's object storage rather than the instance's, `publication_lock` folded 32 bits where it claimed 64, and `ResolvedProfile` had taken `write_profiles`'s doc block. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: a deferring dbt run never publishes the state it read `publishes_ownership` reads the CALLER's overrides, so a descriptor that already narrows `select` needs none and a run of it with `defer: true` published. A deferring run built some of the relations its manifest names and resolved the rest out of the state it read, so recording that manifest claims relations nothing built — and a model renamed since is recorded under a name only a full build creates, breaking every later deferral until one repairs it. Also: `publication_lock` parsed 16 hex digits as `i64`, which overflows for every digest with the top bit set — half of them — collapsing those environments onto one advisory key; and a failure to open the transaction returned without dropping the objects already uploaded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: only a deployed dbt run publishes state, and key its objects per execution A preview carries a caller-supplied `script_hash` into `runnable_id` (`run_preview_script`), so the version guard alone let anyone who may run a job publish arbitrary content as a deployed script's deferral state. The job's KIND is checked beside it now. Verified: a preview submitted with the deployed path and hash builds and leaves the row untouched. Object keys carry a per-execution nonce. Zombie recovery re-runs a job under its own id, so keyed on that alone a second attempt overwrote the objects the first attempt's committed row still named, then read those same keys back as displaced and dropped them — leaving the row unreadable. The displaced set is also filtered against this publication's own keys, so the invariant is stated rather than re-derived from the key format. A project-owned `profiles.yml` that templates its schema or database is refused a deferral: dbt renders those and Windmill does not, so two renderings resolve to one `relation_root` and would share one environment key. Plainly absent is left alone — that is the adapter's default, which does not move. The deferral log line now says the run publishes no state of its own, which was otherwise invisible. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: a templated profile location publishes no dbt state either, on every path A `dbt_profile` resource is one block of the user's own `profiles.yml` copied through unchanged, and `profile.schema` is written as given, so either can carry a template dbt renders and this runtime does not — exactly as a project-owned file can. Only the project-owned path detected it. And the refusal now covers publication as well as deferral: a published template would sit under a key a literal profile shares, so de-templating later would make that stale manifest readable as the new location's. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: recognise Jinja statement blocks as a rendered dbt profile location dbt renders a profile through Jinja, so `{% if env_var('ENV') == 'prod' %}…{% endif %}` moves a schema exactly as an `env_var()` substitution does — and only `{{` was detected, so such a profile published and deferred under one environment key for every rendering. One predicate now serves both profile paths, with a test for each delimiter. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: a dbt state read outruns successive publications rather than one The loader re-read once, which answers a single publication overtaking it: a reader takes no lock and the advisory lock is released before the displaced objects are dropped, so back-to-back publications could each overtake the same read and the second was reported as a missing object. It now re-reads for as long as the row keeps MOVING, bounded, and reports only when an unmoved row's objects are genuinely gone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: a dbt state read outruns successive publications, not one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: name both ways a dbt state read can fail Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: length-prefix the dbt environment key's components A dbt target name and a schema are both the user's own strings, so joining them on `|` let one component spell another tuple's key: `prod|analytics` + `scratch` and `prod` + `analytics|scratch` were one environment, and a profile moving between them read as the same one rather than as one nothing has published — the collision the key exists to prevent. The schema and database are also taken apart now rather than through `relation_root`'s own join, so neither can absorb the other's delimiter. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: name the dbt environment in words where a message shows it The key is length-prefixed for storage, which is not something to put in front of a caller: the "nothing published yet" refusal now reads "warehouse `main`, target `prod`, relations in `dbt_wh_defer.analytics`". The worked example of the encoding also miscounted a component. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: delete a script version in the transaction that cleans up after it `delete_script_by_hash` soft-deleted through the pool, committing before the cleanup that follows it in `tx`. In that window the path has no live version, so a concurrent deploy can take it — and `clear_dbt_script_state_if_path_retired` then finds that new script live, keeps the deleted project's dbt state, and leaves the replacement able to defer through its manifest. The update moves into the same transaction, which is what `archive_script_by_hash` beside it already does. The retirement guard itself was pinned by nothing: the existing test moved the only row away before calling the conditional clear, so it could not fail. `state_goes_only_once_no_live_version_is_left` covers both directions — a second live version keeps the state, the last one leaving takes it — and fails if the predicate is inverted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: archive a script by path in the transaction that cleans up after it The last of the four routes still writing outside its own cleanup transaction. Archived on its own, a cleanup that then fails leaves dbt state at a path no live version occupies, and whatever is created there next can defer through it. The by-hash archive and both deletes already take their write in `tx`; this makes the set uniform. Two comments beside those clears still called the state the RETRY state alone, which the rename made false — they cover both halves now — and the merged verification list had two `11.`, main's #10978 having inserted an item above it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: refuse a dbt state selector the engines resolve inconsistently `state:`, `result:` and `source_status:` selectors resolve against the artifacts in `--state`, which only a deferring run is handed. The engines disagree about what happens without one, and two of the three disagree silently: dbt-core 1.x raises, but dbt-sa-cli 2.x and fusion read a missing state as an empty one and exit 0, so `state:modified` builds nothing and `state:new` builds the whole project, each reporting success. Refuse them up front instead, naming `defer`. From the descriptor they are refused outright, since that selection also decides which nodes the script owns and the deploy resolves it with no state at all. `source_status:` is refused under any setting: it compares `sources.json`, which no run publishes here. A caller's selection is now allowed to match nothing, which is what `state:modified+` returns when nothing changed since the published state. It is stored as that run's own snapshot and never becomes what the script owns, so the ownership-wipe the refusal guarded against cannot happen. The descriptor's selection still may not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: refuse a dbt result selector the published state cannot answer Round 18 findings. Codex P1: `defer` alone was enough to allow a `result:` selector, but a build recovered by node retry publishes a manifest with no `run_results.json` — the only file such a selector reads. dbt-core then raises an internal error and the Rust engines match nothing and exit 0. The deferral now reports whether the state carries results, and a `result:` selection against one that does not is refused, naming the run that published it. Claude P2: a `parse` returns before `defer` is read, so its deferral is always absent and "turn `defer` on" was advice that led nowhere. The check now distinguishes a run that could defer from a command that never does, and the parse path says so. Codex P2 / Claude P2: the roadmap still listed `state:modified` as out of scope while the same file documented it as working. Narrowed both that line and the scope list to the slim-CI work that genuinely remains. Also pins the invariant the relaxed empty-selection guard rests on: an overridden selection must not publish ownership, or an empty caller selection would wipe the script's graph. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: exempt an empty dbt selection by method, not by who chose it Round 19 findings. Codex P1: the empty-selection exemption keyed on whether the caller overrode the selection, so a misspelled model name resolved to nothing, passed the guard and reported a build that did its work. Key it on the selector instead: only a `state:` or `result:` method may match nothing, its empty answer being a real one. Every other selection matching nothing is refused again, from a run as from the descriptor, each with the message that applies to it. Claude P2: the spec still described a node-retry-recovered publication as one where `result:` selectors merely lose their input, which the previous commit stopped being true, and the section stating the selector rules recorded neither the `result:`-without-results refusal nor the `parse` one. Both written down. Also drops the refusal's claim that the publishing run WAS recovered by node retry: an unreadable file reaches the same absent-results state, and the remedy is the same either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: record why an exempted empty dbt selection cannot wipe the graph The safety argument left with the origin-based condition it justified. Under the method-based one it is a consequence of the descriptor refusal in check_state_selectors, two hops from this site, so state it here: relaxing that refusal would let a descriptor-narrowed `state:modified+` reach the exemption and be ingested as owning nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>