feat: capture managed-materialize output schema as asset metadata (#2a) (#9812)

* feat: capture managed-materialize output schema as asset metadata (#2a)

After a managed `// materialize` run, capture the producer's output schema
via a DESCRIBE folded into the existing one-row summary read (no extra
round-trip) and persist it in a new versioned `materialized_asset_schema`
sidecar table. This is the producer-side capture that pipeline parity gap
#2b (save-time consumer-ref contract enforcement) will read back.

- materialized_asset_schema sidecar (asset-level grain), versioned: a new
  version row is inserted only when the captured column set changes.
- output_schema column added to the materialize summary codegen.
- worker extracts + records the schema on a successful materialize.
- /assets/asset_schemas read endpoint exposing the evolution history.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address CI review on schema capture (partition col, order, status gate)

- exclude the synthetic `_wm_partition` column from the captured schema for
  partitioned assets, so the recorded contract is the producer's logical
  output, not Windmill's storage detail (claude/cubic P1).
- make the captured column list explicitly ordered (`row_number()` over the
  DESCRIBE + `list(... ORDER BY)`), so the `list()` aggregate can't reorder
  columns and spuriously bump the schema version (cubic P2).
- gate the API `record_materialization` schema upsert on a `Materialized`
  status, so a failed/running write (or a client attaching a schema to one)
  can't advance the schema history (cubic P2).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address Codex review (manual-mode schema gate + auth contract docs)

- gate output_schema extraction on the managed (`Some((Some(_), _))`) path so a
  `// materialize manual` run — whose result is the user's own query output —
  can't persist a caller-shaped `output_schema` into materialized_asset_schema
  (Codex P2). Verified e2e: a manual run returning a fabricated
  `output_schema:[{injected,EVIL}]` records the partition but writes no schema
  version, while the managed path still captures normally.
- document the authorization contract on the new public `record_asset_schema`
  and `list_asset_schemas` helpers: they perform no access control (mirroring
  the materialized_partition siblings) and require callers to pass a
  workspace-authorized executor (Codex P1).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(frontend): schema-history tab on the ducklake asset node (#2a)

Adds a "Schema" tab to DucklakeAssetPanel surfacing the captured output-schema
versions persisted by the materialize run. Master-detail (mirrors the History
tab): the version list (newest first, newest auto-selected) shows column count +
snapshot + capture time; selecting a version renders its column/type table.
Reads the GET /assets/asset_schemas endpoint via raw fetch, matching the sibling
PartitionStatusGrid convention (these materialization endpoints are not in the
generated client).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat: schema tab is strategy-aware (history vs fixed schema)

Only a whole-table `replace` producer (CREATE OR REPLACE) can change columns
run-to-run; `append`/`merge`/partitioned writes INSERT into a fixed-schema
table, so their schema is pinned at first materialize and the "history" framing
is degenerate (always one version).

- backend: surface the managed `materialize_strategy` (`replace`/`append`/
  `merge`) on the asset-graph runnable node, alongside the existing
  `partition_kind` (same parse-from-annotation path).
- frontend: the pipeline page derives `schemaCanEvolve` for the selected asset
  from its write-producer (`replace` && not partitioned) and threads it to the
  Schema tab. Evolvable → master-detail version history; fixed → a single
  current-schema table with a short "schema is fixed" note. Unknown defaults to
  evolvable so real history is never hidden.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: schemaCanEvolve fails open on unknown producer strategy

Previously a producer present but missing `materialize_strategy` (e.g. a
draft-overlay runnable, synthesized without the field) fell through to
canEvolve=false, hiding captured history behind the fixed-schema view —
contradicting the "unknown defaults to evolvable" intent.

Now the fixed view shows only when *every* producer is a known insert-style
write (append/merge, or partitioned replace); any producer with unknown
(missing) strategy is treated as evolvable, so real history is never hidden.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Ruben Fiszel
2026-06-26 19:46:48 +02:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 44c25de418
commit ade74b297f
19 changed files with 1008 additions and 46 deletions
@@ -0,0 +1,33 @@
{
"db_name": "PostgreSQL",
"query": "UPDATE materialized_asset_schema\n SET snapshot_id = $5, job_id = $6, captured_at = now()\n WHERE workspace_id = $1 AND asset_kind = $2 AND asset_path = $3\n AND version = $4",
"describe": {
"columns": [],
"parameters": {
"Left": [
"Text",
{
"Custom": {
"name": "asset_kind",
"kind": {
"Enum": [
"s3object",
"resource",
"variable",
"ducklake",
"datatable",
"volume"
]
}
}
},
"Text",
"Int8",
"Int8",
"Uuid"
]
},
"nullable": []
},
"hash": "01732ca02b1888145c48c4e51e5b5829657224a743af9c0b2d5a140ad70e13dd"
}
@@ -0,0 +1,44 @@
{
"db_name": "PostgreSQL",
"query": "SELECT version, columns AS \"columns: Json<Vec<SchemaColumn>>\"\n FROM materialized_asset_schema\n WHERE workspace_id = $1 AND asset_kind = $2 AND asset_path = $3\n ORDER BY version DESC\n LIMIT 1",
"describe": {
"columns": [
{
"ordinal": 0,
"name": "version",
"type_info": "Int8"
},
{
"ordinal": 1,
"name": "columns: Json<Vec<SchemaColumn>>",
"type_info": "Jsonb"
}
],
"parameters": {
"Left": [
"Text",
{
"Custom": {
"name": "asset_kind",
"kind": {
"Enum": [
"s3object",
"resource",
"variable",
"ducklake",
"datatable",
"volume"
]
}
}
},
"Text"
]
},
"nullable": [
false,
false
]
},
"hash": "231475dc825518aa88f562698f8061d2562c3ac4af8d088f9eea9f0bc11d5fe7"
}
@@ -0,0 +1,62 @@
{
"db_name": "PostgreSQL",
"query": "SELECT version, columns AS \"columns: Json<Vec<SchemaColumn>>\",\n snapshot_id, job_id, captured_at\n FROM materialized_asset_schema\n WHERE workspace_id = $1 AND asset_kind = $2 AND asset_path = $3\n ORDER BY version DESC",
"describe": {
"columns": [
{
"ordinal": 0,
"name": "version",
"type_info": "Int8"
},
{
"ordinal": 1,
"name": "columns: Json<Vec<SchemaColumn>>",
"type_info": "Jsonb"
},
{
"ordinal": 2,
"name": "snapshot_id",
"type_info": "Int8"
},
{
"ordinal": 3,
"name": "job_id",
"type_info": "Uuid"
},
{
"ordinal": 4,
"name": "captured_at",
"type_info": "Timestamptz"
}
],
"parameters": {
"Left": [
"Text",
{
"Custom": {
"name": "asset_kind",
"kind": {
"Enum": [
"s3object",
"resource",
"variable",
"ducklake",
"datatable",
"volume"
]
}
}
},
"Text"
]
},
"nullable": [
false,
false,
true,
true,
false
]
},
"hash": "2a28f13c1f06a77bc9a164393cccf0d8d2a3d7e8552d8ad09512d2d6bf27c3fb"
}
@@ -0,0 +1,34 @@
{
"db_name": "PostgreSQL",
"query": "INSERT INTO materialized_asset_schema\n (workspace_id, asset_kind, asset_path, version, columns,\n snapshot_id, job_id, captured_at)\n VALUES ($1, $2, $3, $4, $5, $6, $7, now())",
"describe": {
"columns": [],
"parameters": {
"Left": [
"Varchar",
{
"Custom": {
"name": "asset_kind",
"kind": {
"Enum": [
"s3object",
"resource",
"variable",
"ducklake",
"datatable",
"volume"
]
}
}
},
"Varchar",
"Int8",
"Jsonb",
"Int8",
"Uuid"
]
},
"nullable": []
},
"hash": "8e829a7358ea47100c99e78067266eb7a83e312b88f0244d1c517b03a9bcdea0"
}
@@ -0,0 +1,22 @@
{
"db_name": "PostgreSQL",
"query": "SELECT pg_advisory_xact_lock(hashtextextended($1, 0::int8))",
"describe": {
"columns": [
{
"ordinal": 0,
"name": "pg_advisory_xact_lock",
"type_info": "Void"
}
],
"parameters": {
"Left": [
"Text"
]
},
"nullable": [
null
]
},
"hash": "b876b16b660bd1c3799abadc2c16d0f58f49a45647525658acb909a494877238"
}