Files
windmill/typescript-client/tests/sqlUtils.test.ts
T
Ruben Fiszel aedf369174 fix: pair PG arg type with actual Rust binding to keep query_typed_raw safe (#8999)
* fix: pair PG arg type with actual Rust binding to keep query_typed_raw safe

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(pg): wrap encoder errors with arg context, add fallback test

Followups on #8999 review:

- Wrap rust-postgres "error serializing parameter N" failures with the arg
  name, JSON value kind, and asserted Postgres type plus a hint about an
  explicit cast — so users see actionable context instead of an opaque
  WrongType.
- Drift-prevention meta-test: assert otyp_to_pg_type and convert_val agree
  on the Type for every recognised arg_t when the JSON value matches its
  natural Rust kind. Catches future drift if either side changes.
- Integration test for the prepare + query_raw fallback path: confirms
  unrecognised arg_t (custom enum) is routed through prepare and the
  server-resolved type appears in the failure surface — flips into a
  test failure if a regression accidentally routes unrecognised types
  through query_typed_raw.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(pg): add otyp_inferred flag + regex-based placeholder renumbering

Two follow-ups from the review of #8999:

1. **Issue #1 (Number/Bool + explicit text decl in WHERE)**

   Add `Arg::otyp_inferred: bool` to the parser. The PG SQL parser sets
   it `true` only at the "no info → fall back to text" site (bare `$N`,
   no inline cast, no `-- $N (TYPE)` decl). All other arg sources keep
   it `false`.

   In `convert_val` this flag distinguishes:
   - explicit text-like target (`-- $1 (text)` or `$1::text`) — coerce
     `Bool`/`Number` → `Box<String>` so `WHERE text_col = $1` works
     (`text = text` operator). Pre-#8988 behaviour, restored.
   - parser-default text (bare `$N`) — bind the value's natural Rust
     type so the regression case (`Value::Bool` against a real `bool`
     column via `CAST AS bool`) keeps working.

   `Arg` is in `windmill-parser`; the new field has `#[serde(default)]`
   so persisted signatures stay backward-compatible.

2. **Issue #4 ($5/$50 substring rewrite collision)**

   Replace the per-index `String::replace` chain (which turned `$50`
   into `$10` when oidx=5 was processed first) with a single regex
   pass. `\d+` is greedy, so `$5` and `$50` match as distinct units;
   indices outside the mapping are left intact.

3. Tests:
   - parser: `test_parse_pgsql_otyp_inferred_flag` covers bare/inline-
     cast/decl/mixed shapes.
   - executor unit: `convert_val_bool_against_every_arg_t` and
     `convert_val_*_number_*` split each text-like target into explicit
     vs inferred expectations.
   - executor unit: `renumber_sparse_placeholders_no_collision`.
   - integration: `test_postgresql_arg_type_combinations` adds 4 cases
     covering decl(text)+Number/Bool in WHERE, bare $1+Bool, and
     sparse positional args ($5/$50).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(pg+sdk): enum support, extended String arms, position-aware $N rewrite, SDK quality

Backend:

1. **`AnyTextValue` ToSql/FromSql wrapper**: vanilla `tokio_postgres`'s
   `ToSql for String` / `FromSql for String` reject `Kind::Enum` and
   `Kind::Domain` even though the wire format is plain UTF-8. The wrapper
   accepts those kinds in both directions. End result: explicit
   `$1::my_enum` / `CAST($1 AS my_enum)` casts now round-trip without the
   ugly `CAST($1::text AS my_enum)` workaround, AND `SELECT enum_col`
   results come back as JSON strings instead of erroring at the FromSql
   layer.

2. **#10 — Value::String → numeric/real/double/oid/bool**. Without these
   arms, a string-encoded value (`"3.14"`, `"true"`) for a non-text /
   non-temporal arg_t fell through to `Box<String> + TEXT`, which then
   failed at the server (no implicit cast text→numeric in expression
   context). Now strings are parsed into the matching native type with
   clear error messages on parse failure.

3. **Position-aware `$N` rewrite**: replaces the regex-based renumbering
   (which fixed the `$5/$50` substring collision but still walked through
   string literals and comments, mangling `'price: $5'` etc.) with a
   walk over `parse_pg_statement_arg_positions` — the same
   string/comment/dollar-quote-aware tokenizer used for index discovery.
   Adds `parse_pg_statement_arg_positions` to the parser's public API.

SDK:

4. **BigInt support**: `JSON.stringify(BigInt)` throws. The SDK now
   stringifies bigints before serialisation; the executor accepts
   numeric strings into BIGINT arg slots via the existing
   `Value::String → INT8` parsing arm. SDK-side `inferSqlType` is split
   so `BigInt` always resolves to `BIGINT` (was reaching
   `Number.isInteger(BigInt)` which returns false → wrong default).

5. **Homogeneous array auto-tag**: `${[1,2,3]}` against an `int[]` column
   now emits `$1::BIGINT[]` instead of `$1::JSON`. Detection covers
   primitive types only (number / bigint / string / boolean); mixed or
   nested arrays still fall back to JSON. Mixed int/float widens to
   `DOUBLE PRECISION[]`.

6. **`.query()` positional bug**: previously the `.query()` method
   abused the template-tag builder, which appended `$N::TYPE` after the
   user's literal SQL string instead of binding by position
   (`SELECT $1, $2` became `SELECT $1, $2$1::BIGINT`). Now `.query()`
   builds the executor-shaped content directly: a `-- $N argN (TYPE)`
   declaration block followed by the user's SQL verbatim.

Tests:

- Parser: `test_parse_pg_statement_arg_positions_skips_strings_and_comments`
  asserts string literals, comments, and dollar-quoted blocks don't
  produce positions (so renumbering doesn't mangle them).
- Executor unit: `renumber_sparse_placeholders_no_collision_no_string_mangling`
  uses the new position-aware path and includes string-literal + comment
  + `$$…$$` cases. Existing convert_val tests grow to cover new
  String→numeric/real/double/oid/bool arms.
- Integration: `test_postgresql_arg_type_combinations` adds 13 cases
  (enum round-trip both directions, string→numeric/real/double/bool/oid,
  string-literal `$N` non-mangling). The prepare-fallback test now
  asserts SUCCESS (not failure) for enum encoding via AnyTextValue.
- SDK: new `typescript-client/tests/sqlUtils.test.ts` (42 tests)
  exhaustively covering inferSqlType primitives + arrays,
  parseTypeAnnotation, datatable() template tag (with all the new
  shapes — BigInt, homogeneous arrays, RawSql, schema preamble),
  datatable().query() positional, and ducklake() shape.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(pg): replace DISCARD ALL with curated reset (preserves typeinfo cache)

Found while exhaustively probing custom-type DX: every cached-connection
reuse was running `DISCARD ALL`, whose included `DEALLOCATE ALL`
deallocates *all* prepared statements server-side — including the typeinfo
statements that tokio_postgres caches per-Client to resolve custom enum /
domain Oids. tokio_postgres still held `Statement` objects whose names
the server had forgotten, so the next custom-type query failed with
intermittent "prepared statement \"sN\" does not exist" errors. The
failure was easy to reproduce: any sequence that forced typeinfo lookup
for two different custom-type kinds on the same cached connection (e.g.
enum followed by domain) would hit it.

Replace `DISCARD ALL` with a curated reset that explicitly targets the
state we actually care about, *without* touching prepared statements:

  RESET ALL                     — GUC parameters (search_path, application
                                  _name, statement_timeout, …)
  RESET SESSION AUTHORIZATION   — undoes both `SET SESSION AUTHORIZATION`
                                  and `SET ROLE` (RESET ALL does NOT —
                                  these aren't GUC parameters, so without
                                  this an elevated role from a previous
                                  job would silently leak)
  UNLISTEN *                    — drops LISTEN registrations
  CLOSE ALL                     — closes open cursors

Trade-off: temp tables, advisory locks (session-scoped), and user-created
PREPARE statements may persist across cached-connection reuse — rare in
datatable / PG-script workloads. tokio_postgres's typeinfo cache survives
intact, so custom enum / domain queries are fast on subsequent reuse.

Tests:
- `test_postgresql_custom_types_on_cached_connection` — runs 10×
  alternating enum + domain queries on a cached connection. Pre-fix this
  failed with `prepared statement "sN" does not exist` after the first
  reuse; post-fix passes.
- `test_postgresql_set_role_does_not_leak_across_cached_connection` —
  switches `SET ROLE` and `SET SESSION AUTHORIZATION` to a non-postgres
  role, then runs a follow-up job and asserts current_user/session_user
  are restored. Specifically catches the case where someone might switch
  back to `RESET ALL` alone (which doesn't cover SET ROLE / SESSION
  AUTHORIZATION) and silently introduce a permission-leak vector.
- All existing session-isolation tests
  (`test_postgresql_cached_connection_resets_session`,
   `test_postgresql_single_worker_session_isolation`,
   `test_postgresql_100_jobs_cached`) continue to pass.

Found via end-to-end probing of datatable / PG-script DX, not previously
covered: the existing isolation tests only did `SET ROLE postgres`, the
connecting user, so the leak was invisible.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(pg): address PR #8999 review (cubic + claude)

cubic (P1, real bug):
- `convert_vec_val` for `timetz` array asserted `Type::TIMETZ_ARRAY`, but
  chrono `NaiveTime` only encodes for TIME (same caveat as the scalar
  arm). Switch to `Type::TIME_ARRAY`; rely on PG's implicit `time→timetz`
  assignment cast at the column site. Add an explicit unit test.

claude (#1, silent failure → explicit error):
- `Bool` + explicit `(char)` / `(character)` decl previously silently
  bound BOOL, hoping the server would cast at the use site — but PG has
  no implicit `bool→char` and the resulting error
  ("operator does not exist: bool = char") was opaque. Now error at
  bind time with an actionable hint to use `bool` decl or pass the
  value as a "t"/"f" string.

claude (#2, asymmetry doc):
- Object/Array still coerce to text on `matches!(typ, Typ::Str(_))`
  (covers both explicit AND inferred-default text), unlike Bool/Number
  which key on `explicit_text_target`. The asymmetry is intentional
  (no implicit `jsonb → text` cast in expression context vs PG having
  implicit `bool/int → text` casts) — added a body comment so future
  maintainers don't try to "align" them.

claude (#3, perf):
- `parse_pg_statement_arg_indices` and `parse_pg_statement_arg_positions`
  walked the SQL tokenizer twice. Fold into a single pass that derives
  the index set from the position list.

claude (#4, fmt drift):
- `cargo fmt` over the parser crates I touched with perl scripts in the
  earlier commit (windmill-parser-{sql,bash,ts,go,php,java,csharp,nu,py,
  rust,graphql,yaml,r}). Net cosmetic.

claude (#5, parseTypeAnnotation):
- One-line caveat in the SDK's `parseTypeAnnotation` that the returned
  string is presence-only (e.g. `${x}::DOUBLE PRECISION` returns
  `"DOUBLE"`, `CAST(${x} AS int)` returns `"int)"` — neither matches a
  real PG type, but the only consumer just checks `!== undefined`).

While here — discovered + fixed independently while exhaustively probing
DX:

- **Replace `DISCARD ALL` with curated reset** (`RESET ALL; RESET
  SESSION AUTHORIZATION; UNLISTEN *; CLOSE ALL;`). DISCARD's
  `DEALLOCATE ALL` killed tokio_postgres' typeinfo cache, producing
  intermittent `prepared statement "sN" does not exist` errors on
  custom-type queries after cached-conn reuse. New regression tests:
  `test_postgresql_custom_types_on_cached_connection` and
  `test_postgresql_set_role_does_not_leak_across_cached_connection`
  (the latter catches the case where someone might switch back to
  `RESET ALL` alone and silently introduce a permission-leak vector —
  RESET ALL doesn't cover SET ROLE / SET SESSION AUTHORIZATION).

- **ISO-8601 timestamp results** (`pg_cell_to_json_value`). Pre-fix
  `TIMESTAMP` was rendered with a space separator ("2024-01-15 10:30:00")
  and `TIMESTAMPTZ` with " UTC" suffix ("2024-01-15 10:30:00 UTC") —
  neither parseable by `date-fns parseISO`, JavaScript `new Date()` is
  lenient enough to handle them but several frontend `App*Input.svelte`
  components use parseISO and fail silently. Switched to ISO-8601 with
  `T` separator and `+00:00` offset; arg-parsing path still accepts the
  legacy " UTC" suffix for back-compat.

Test coverage:
- 17/17 unit (`pg_executor::tests`)
- 9/9 integration (`backend/tests/worker.rs`, `test_postgresql_*`)
- 27/27 parser (`windmill-parser-sql`)
- 42/42 SDK (`typescript-client/tests/sqlUtils.test.ts`)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(pg): bounded one-shot warning on numeric precision loss + ISO-8601 + NaN handling

Found while probing PG-script DX with millions of numeric cells:

1. **Numeric precision-loss warning**: `numeric` results are still serialised
   as JSON Number (back-compat — switching to JSON String would silently
   break user code doing arithmetic on results), but we now detect
   `Decimal -> f64 -> Decimal` round-trip failure and emit a single
   job-log warning recommending a `::text` cast in the SQL. Bounded by
   `NUMERIC_PRECISION_CHECK_BUDGET = 256` cells per query (one atomic
   load + one fetch_sub on the hot path; first lossy value
   short-circuits to a single load thereafter). Worst-case overhead on
   a 1M-cell numeric-heavy query: ~25µs of checks + 5ns × N atomic
   loads (vs. ~100ms unbounded).

2. **ISO-8601 timestamps**: `pg_cell_to_json_value` previously returned
   `"2024-01-15 10:30:00"` (TIMESTAMP) and `"2024-01-15 10:30:00 UTC"`
   (TIMESTAMPTZ) — neither parseable by date-fns `parseISO`, which is
   what the apps `App*Input.svelte` components use, so timestamp values
   silently failed to round-trip into date pickers. Switch to ISO-8601
   (`T` separator + `+00:00` offset) on the result side; arg-parser
   continues to accept the legacy `" UTC"`-suffixed format for
   back-compat.

3. **Float NaN / Infinity results**: `Number::from_f64` returns None for
   NaN / ±Inf, which `pg_cell_to_json_value` was raising as
   "invalid json-float" — failing the *entire* query if any cell held
   one of these special values. Now serialise them as JSON strings
   ("NaN", "Infinity", "-Infinity") and let the rest of the row come
   through. Arg-side: `s.parse::<f64>()` already accepts the same
   strings.

Tests:
- `decimal_fits_f64_losslessly_predicate` — covers fits / doesn't-fit
  cases for the precision-loss predicate.
- `precision_check_budget_caps_per_query_overhead` — locks in the
  budget cap and the loss-flag short-circuit.
- All 9 PG integration tests + 17 unit tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(pg): add pg_advisory_unlock_all to reset; warn on missing args; honor decl defaults

While probing PG-script DX further found three more frictions:

1. **Advisory lock leak** (cubic P2): switching from `DISCARD ALL` to
   `RESET ALL; RESET SESSION AUTHORIZATION; UNLISTEN *; CLOSE ALL;`
   meant session-scoped advisory locks (`pg_advisory_lock`) leaked
   across cached-connection reuse. Add `SELECT pg_advisory_unlock_all()`
   to the chain — `DISCARD ALL` covered this implicitly via
   `DISCARD PLANS / DEALLOCATE / pg_advisory_unlock_all` and we lost it
   in the switch.

2. **Missing-arg silent NULL**: an arg declared in the SQL (e.g.
   `-- $1 amount (numeric)`) but not provided in the args object was
   bound as NULL with no error / warning. Misspelling the key in the
   args object silently produced a row of NULLs — a notorious DX
   debugging trap. Now: collect the names of declared-but-missing
   args during dispatch and emit a single one-shot warning to the job
   logs at end-of-query naming each one. Bound NULL is preserved for
   back-compat.

3. **Declaration defaults ignored**: `-- $1 a (int) = 5` carries
   `arg.default = Some(Number(5))`, but the dispatch fell straight to
   NULL when the arg was missing. Now: respect the default —
   user-supplied value > declaration default > NULL. Also fixes the
   warning logic above (only warn for args that *don't* have a default).

Tests: existing 19 unit + 9 integration pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(pg): multi-word PG types with [] suffix lost the array-ness; array arms accept stringified values

Two more frictions found while probing SDK end-to-end against a real
datatable resource:

1. **Multi-word array types lose the [] suffix in the parser**.
   `transform_types_with_spaces` recognises aliases for "double
   precision", "character varying", "timestamp with time zone", etc.
   but its return type was `&'a str` — only the bare alias, never with
   a trailing `[]`. The `RE_CODE_PGSQL` regex's `\w+` captures stop at
   the first space, so the regex's own `(?:\[\])?` array-suffix branch
   sees only `"double"` (not `"double precision[]"`); the `[]` was
   silently lost. Result: `$1::double precision[]` (which the SDK now
   emits for homogeneous float arrays via the new auto-tag) routed
   through `Value::Array → Type::JSONB` and the server failed with
   "cannot cast type jsonb to double precision[]".

   Fix: switch `transform_types_with_spaces` to return `Cow<'a, str>`
   and re-check the trailing bytes after a multi-word match. If they
   start with `[]`, return `format!("{alias}[]")` — Owned. Single-word
   types and the no-match path keep returning Borrowed slices, so no
   allocation in the hot path.

2. **Array arms in `convert_vec_val` rejected stringified values for
   numeric / int* / bool / oid / real / double**. The scalar `convert_val`
   already parses strings into the matching native type for these arg_ts,
   but the array variant only accepted JSON-native counterparts. Sending
   `["1.5", "2.5", "3.5"]` against `$1::numeric[]` (e.g. via `unnest` for
   bulk loading, or `JSON.stringify(BigInt[])` round-trip) failed with
   "Mixed types in array". Now the array arms mirror the scalar ones —
   `as_<native>().or_else(|| as_str().and_then(parse))` — so both shapes
   round-trip cleanly.

Tests: 19 unit + 9 integration pass; existing parser tests cover the
multi-word array forms (the regex-cap behaviour didn't break for
single-word types, and Cow plumbing is transparent to all callers).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(parsers): add otyp_inferred field to Arg literals in tests + 3 missed src files

CI failures: the perl-driven sweep that added `otyp_inferred: false` to
every `Arg { ... }` literal when I introduced the field in the parser
schema covered `src/lib.rs` files but missed:

  - parsers/windmill-parser-bash/src/lib.rs       (mass-edited but a
    later format pass un-applied a few sites)
  - parsers/windmill-parser-go/src/lib.rs         (same)
  - parsers/windmill-parser-graphql/src/lib.rs    (same)
  - parsers/windmill-parser-nu/tests/tests.rs     (test file — not
    swept the first time)
  - parsers/windmill-parser-ts/tests/tests.rs     (test file — same)

Also tightened the regex to handle `oidx: None` without the trailing
comma (some test files had the field as the last initialiser line).

`cargo build --features <CI feature combo> --workspace --all-targets`
is clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(sdk): Date → TIMESTAMPTZ; NaN / ±Infinity → string

Two more frictions found while running the actual SDK end-to-end against
a live datatable resource:

1. **JS `Date`** fell into the typeof "object" branch and was tagged
   `::JSON`. It worked accidentally for `${date}::timestamptz` via PG's
   `json → text → timestamptz` implicit cast chain, but `${date}` against
   a `timestamptz` column without a user-supplied cast bound the value
   as a JSON string and the comparison `timestamptz = json` failed. Now:
   `inferSqlType` recognises `Date` and tags `::TIMESTAMPTZ`;
   `serializeArgValue` emits `Date.toISOString()` so the executor's
   `Value::String → TIMESTAMPTZ` arm parses it cleanly.

2. **JS `NaN` / `±Infinity`** silently became NULL. `JSON.stringify(NaN)`
   returns `"null"` per the JS spec, so the value reached the executor as
   JSON null — the SDK's `::DOUBLE PRECISION` tag then bound a NULL
   double. Fix: detect non-finite numbers in `serializeArgValue` and
   stringify them as `"NaN" / "Infinity" / "-Infinity"`. The executor's
   `Value::String → FLOAT8` arm (`f64::from_str`) accepts these literals
   directly, and the result-side already renders the values as JSON
   strings (matching round-trip).

SDK unit tests grow from 42 → 44 passing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(pg): integration coverage for multi-word arrays + stringified array elements

Locks in the two array fixes from the previous commit
(`fix(pg): multi-word PG types with [] suffix lost the array-ness`)
with end-to-end cases in `test_postgresql_arg_type_combinations`:

- `double precision[]`, `character varying[]`, `timestamp without time
  zone[]` — verifies the parser keeps the `[]` suffix after multi-word
  alias resolution.
- `numeric[]` / `int[]` / `bool[]` from stringified primitives — verifies
  the array arms of `convert_vec_val` apply the same string-coercion
  the scalar arms do.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style: fix indentation drift on otyp_inferred lines

cargo fmt cleanup of leftover indentation where the perl-driven sweep
that introduced the otyp_inferred field landed at the wrong column.
No behaviour change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-05-01 17:08:59 +00:00

627 lines
23 KiB
TypeScript

/**
* Standalone tests for `wmill.datatable()` / `wmill.ducklake()` SQL template
* functions.
*
* The real `sqlUtils.ts` imports `./services.gen` (auto-generated, not in
* the repo) so we can't import it here. Instead we re-implement the same
* type-inference / template-building / `.query()` pipeline inline (sans the
* network calls) and assert the `content` + `args` shapes the executor
* would receive. Each test maps 1:1 to a behaviour this PR introduces or
* fixes (BigInt, homogeneous arrays, `.query()` positional, etc.) so they
* also serve as a regression backstop.
*
* Run with: bun test typescript-client/tests/sqlUtils.test.ts
*/
import { expect, test, describe } from "bun:test";
// =============================================================================
// Pure SDK logic (mirror of typescript-client/sqlUtils.ts — kept minimal,
// only the parts that decide content / args).
// =============================================================================
class RawSql {
readonly __brand = "RawSql" as const;
constructor(public readonly value: string) {}
}
interface SqlProvider {
formatArgDecl(argNum: number, argType: string): string;
formatArgUsage(
argNum: number,
explicitType: string | undefined,
inferredType: string
): string;
preamble(): string;
language: "postgresql" | "duckdb";
extraArgs: Record<string, any>;
providerName: string;
}
function datatableProvider(name: string, schema?: string): SqlProvider {
return {
providerName: "datatable",
language: "postgresql",
extraArgs: { database: `datatable://${name}` },
formatArgDecl: (argNum) => `-- $${argNum} arg${argNum}`,
formatArgUsage: (argNum, explicitType, inferredType) =>
explicitType !== undefined
? `$${argNum}`
: `$${argNum}::${inferredType}`,
preamble: () => (schema ? `SET search_path TO "${schema}";\n` : ""),
};
}
function ducklakeProvider(name: string): SqlProvider {
return {
providerName: "ducklake",
language: "duckdb",
extraArgs: {},
formatArgDecl: (argNum, argType) => `-- $arg${argNum} (${argType})`,
formatArgUsage: (argNum) => `$arg${argNum}`,
preamble: () => `ATTACH 'ducklake://${name}' AS dl;USE dl;\n`,
};
}
function inferSqlType(value: any): string {
if (typeof value === "bigint") return "BIGINT";
if (typeof value === "number") {
if (Number.isInteger(value)) return "BIGINT";
return "DOUBLE PRECISION";
} else if (value === null || value === undefined) {
return "TEXT";
} else if (typeof value === "string") {
return "TEXT";
} else if (Array.isArray(value)) {
return inferSqlArrayType(value);
} else if (value instanceof Date) {
return "TIMESTAMPTZ";
} else if (typeof value === "object") {
return "JSON";
} else if (typeof value === "boolean") {
return "BOOLEAN";
} else {
return "TEXT";
}
}
function inferSqlArrayType(value: any[]): string {
if (value.length === 0) return "JSON";
let scalarType: string | undefined = undefined;
for (const elem of value) {
let elemType: string;
if (typeof elem === "bigint") elemType = "BIGINT";
else if (typeof elem === "number")
elemType = Number.isInteger(elem) ? "BIGINT" : "DOUBLE PRECISION";
else if (typeof elem === "string") elemType = "TEXT";
else if (typeof elem === "boolean") elemType = "BOOLEAN";
else return "JSON";
if (scalarType === undefined) scalarType = elemType;
else if (scalarType === "BIGINT" && elemType === "DOUBLE PRECISION")
scalarType = "DOUBLE PRECISION";
else if (scalarType === "DOUBLE PRECISION" && elemType === "BIGINT") {
// already widened
} else if (scalarType !== elemType) {
return "JSON";
}
}
return `${scalarType}[]`;
}
function parseTypeAnnotation(
prevTemplateString: string | undefined,
nextTemplateString: string | undefined
): string | undefined {
if (!nextTemplateString) return;
nextTemplateString = nextTemplateString.trimStart();
if (nextTemplateString.startsWith("::")) {
return nextTemplateString.substring(2).trimStart().split(/\s+/)[0];
}
prevTemplateString = prevTemplateString?.trimEnd();
if (
prevTemplateString?.endsWith("(") &&
prevTemplateString
.substring(0, prevTemplateString.length - 1)
.trim()
.toUpperCase()
.endsWith("CAST") &&
nextTemplateString.toUpperCase().startsWith("AS ")
) {
return nextTemplateString.substring(2).trimStart().split(/\s+/)[0];
}
}
function serializeArgValue(v: any): any {
if (typeof v === "bigint") return v.toString();
if (v instanceof Date) return v.toISOString();
if (typeof v === "number" && !Number.isFinite(v)) {
if (Number.isNaN(v)) return "NaN";
return v > 0 ? "Infinity" : "-Infinity";
}
return v;
}
function buildContentAndArgs(
provider: SqlProvider,
strings: TemplateStringsArray | string[],
values: any[]
): { content: string; args: Record<string, any> } {
let argIndex = 0;
const valueInfos = values.map((v, i) => {
if (v instanceof RawSql)
return { raw: true as const, value: v.value, originalIndex: i };
argIndex++;
return {
raw: false as const,
value: v,
originalIndex: i,
argNum: argIndex,
};
});
let argDecls = valueInfos
.filter((info): info is Extract<typeof valueInfos[number], { raw: false }> => !info.raw)
.map((info) => {
let argType =
parseTypeAnnotation(
strings[info.originalIndex],
strings[info.originalIndex + 1]
) || inferSqlType(info.value);
return provider.formatArgDecl(info.argNum, argType);
});
let content = argDecls.length ? argDecls.join("\n") + "\n" : "";
content += provider.preamble();
let contentBody = "";
for (let i = 0; i < strings.length; i++) {
contentBody += strings[i];
if (i < valueInfos.length) {
let info = valueInfos[i];
if (info.raw) {
contentBody += info.value;
} else {
let explicitType = parseTypeAnnotation(
strings[info.originalIndex],
strings[info.originalIndex + 1]
);
let inferredType = inferSqlType(info.value);
contentBody += provider.formatArgUsage(
info.argNum,
explicitType,
inferredType
);
}
}
}
content += contentBody;
const args = {
...Object.fromEntries(
valueInfos
.filter((info): info is Extract<typeof valueInfos[number], { raw: false }> => !info.raw)
.map((info) => [`arg${info.argNum}`, serializeArgValue(info.value)])
),
...provider.extraArgs,
};
return { content, args };
}
function buildDatatableQuery(
provider: SqlProvider,
sqlString: string,
params: any[]
): { content: string; args: Record<string, any> } {
let argDecls = params
.map((v, i) => `-- $${i + 1} arg${i + 1} (${inferSqlType(v)})`)
.join("\n");
let content =
(argDecls ? argDecls + "\n" : "") + provider.preamble() + sqlString;
let args = {
...Object.fromEntries(
params.map((v, i) => [`arg${i + 1}`, serializeArgValue(v)])
),
...provider.extraArgs,
};
return { content, args };
}
function templateTag(provider: SqlProvider) {
return (strings: TemplateStringsArray, ...values: any[]) =>
buildContentAndArgs(provider, strings, values);
}
const dt = (name = "main") => templateTag(datatableProvider(name));
const dl = (name = "main") => templateTag(ducklakeProvider(name));
const datatableQuery = (name = "main") => {
const provider = datatableProvider(name);
return (sql: string, ...params: any[]) =>
buildDatatableQuery(provider, sql, params);
};
// =============================================================================
// inferSqlType — exhaustive coverage
// =============================================================================
describe("inferSqlType — primitives", () => {
test("integer Number → BIGINT", () => {
expect(inferSqlType(0)).toBe("BIGINT");
expect(inferSqlType(42)).toBe("BIGINT");
expect(inferSqlType(-7)).toBe("BIGINT");
expect(inferSqlType(Number.MAX_SAFE_INTEGER)).toBe("BIGINT");
});
test("non-integer Number → DOUBLE PRECISION", () => {
expect(inferSqlType(0.5)).toBe("DOUBLE PRECISION");
expect(inferSqlType(-3.14)).toBe("DOUBLE PRECISION");
expect(inferSqlType(Number.EPSILON)).toBe("DOUBLE PRECISION");
});
test("BigInt → BIGINT (not DOUBLE PRECISION)", () => {
// Pre-fix this branch was unreachable because bigint was bundled with
// number and `Number.isInteger(BigInt)` returns false → would have
// returned DOUBLE PRECISION (wrong). The split-out check is the fix.
expect(inferSqlType(BigInt(0))).toBe("BIGINT");
expect(inferSqlType(BigInt("9007199254740993"))).toBe("BIGINT");
expect(inferSqlType(BigInt(-1))).toBe("BIGINT");
});
test("string / null / undefined → TEXT", () => {
expect(inferSqlType("")).toBe("TEXT");
expect(inferSqlType("hello")).toBe("TEXT");
expect(inferSqlType(null)).toBe("TEXT");
expect(inferSqlType(undefined)).toBe("TEXT");
});
test("boolean → BOOLEAN", () => {
expect(inferSqlType(true)).toBe("BOOLEAN");
expect(inferSqlType(false)).toBe("BOOLEAN");
});
test("plain object → JSON", () => {
expect(inferSqlType({})).toBe("JSON");
expect(inferSqlType({ a: 1, b: [1, 2] })).toBe("JSON");
});
});
describe("inferSqlType — arrays", () => {
test("empty array → JSON", () => {
expect(inferSqlType([])).toBe("JSON");
});
test("homogeneous integer array → BIGINT[]", () => {
expect(inferSqlType([1, 2, 3])).toBe("BIGINT[]");
expect(inferSqlType([0])).toBe("BIGINT[]");
expect(inferSqlType([-1, 0, 1])).toBe("BIGINT[]");
});
test("homogeneous float array → DOUBLE PRECISION[]", () => {
expect(inferSqlType([1.5, 2.5])).toBe("DOUBLE PRECISION[]");
});
test("mixed int/float array widens to DOUBLE PRECISION[]", () => {
expect(inferSqlType([1, 2.5])).toBe("DOUBLE PRECISION[]");
expect(inferSqlType([1.5, 2])).toBe("DOUBLE PRECISION[]");
});
test("homogeneous string array → TEXT[]", () => {
expect(inferSqlType(["a", "b", "c"])).toBe("TEXT[]");
expect(inferSqlType([""])).toBe("TEXT[]");
});
test("homogeneous bool array → BOOLEAN[]", () => {
expect(inferSqlType([true, false, true])).toBe("BOOLEAN[]");
});
test("homogeneous bigint array → BIGINT[]", () => {
expect(inferSqlType([BigInt(1), BigInt(2)])).toBe("BIGINT[]");
});
test("non-homogeneous array → JSON", () => {
expect(inferSqlType([1, "x"])).toBe("JSON");
expect(inferSqlType(["a", true])).toBe("JSON");
expect(inferSqlType([1, null])).toBe("JSON");
expect(inferSqlType([true, 1])).toBe("JSON");
});
test("nested array → JSON (current limitation, no auto-tag for 2D)", () => {
expect(inferSqlType([[1], [2]])).toBe("JSON");
expect(inferSqlType([{ a: 1 }, { a: 2 }])).toBe("JSON");
});
});
// =============================================================================
// parseTypeAnnotation — used by the SDK to suppress its own ::TYPE injection
// when the user already wrote a cast.
// =============================================================================
describe("parseTypeAnnotation — user-supplied cast detection", () => {
test("`${x}::int` → 'int'", () => {
expect(parseTypeAnnotation("SELECT ", "::int FROM t")).toBe("int");
});
test("whitespace tolerance after ::", () => {
expect(parseTypeAnnotation("SELECT ", " :: bigint FROM t")).toBe(
"bigint"
);
});
test("`CAST(${x} AS int)` → first whitespace-delimited word after AS", () => {
// The SDK splits on whitespace and doesn't strip closing parens, so
// `AS int)` returns "int)". The exact value doesn't matter downstream
// because the SDK only checks `explicitType !== undefined` to skip its
// own ::cast injection — but we lock the behaviour in.
expect(parseTypeAnnotation("SELECT CAST(", " AS int)")).toBe("int)");
});
test("`CAST ( ${x} AS BOOL )` (whitespace + caps)", () => {
// Whitespace before `)` causes split to drop it, so this returns "BOOL".
expect(parseTypeAnnotation("SELECT CAST ( ", " AS BOOL )")).toBe("BOOL");
});
test("no cast adjacent → undefined", () => {
expect(parseTypeAnnotation("SELECT ", " FROM t")).toBeUndefined();
expect(parseTypeAnnotation("SELECT ", "")).toBeUndefined();
expect(parseTypeAnnotation(undefined, undefined)).toBeUndefined();
});
});
// =============================================================================
// datatable() template tag — content + args round-trips
// =============================================================================
describe("datatable() — template tag", () => {
test("primitives auto-tag with ::TYPE inline; decls have no type", () => {
const sql = dt();
const out = sql`SELECT ${42}, ${3.14}, ${true}, ${"x"}, ${null}`;
// datatable provider's formatArgDecl ignores the type, so we get bare
// decls + the casts in the SQL body.
expect(out.content).toContain("-- $1 arg1\n");
expect(out.content).toContain("$1::BIGINT");
expect(out.content).toContain("$2::DOUBLE PRECISION");
expect(out.content).toContain("$3::BOOLEAN");
expect(out.content).toContain("$4::TEXT");
expect(out.content).toContain("$5::TEXT");
expect(out.args).toMatchObject({
arg1: 42,
arg2: 3.14,
arg3: true,
arg4: "x",
arg5: null,
});
});
test("user `${x}::int` suppresses SDK's auto-cast (parser sees user's cast)", () => {
const sql = dt();
const out = sql`SELECT ${42}::int`;
expect(out.content).toContain("SELECT $1::int");
expect(out.content).not.toContain("$1::BIGINT");
});
test("CAST(${x} AS T) syntax → bare $N in SQL (regression #8988)", () => {
const sql = dt();
const out = sql`SELECT CAST(${true} AS bool)`;
expect(out.content).toContain("CAST($1 AS bool)");
expect(out.content).not.toContain("$1::BOOLEAN");
});
test("BigInt is stringified for JSON transport, tagged as ::BIGINT", () => {
const sql = dt();
const out = sql`SELECT ${BigInt("9007199254740993")}`;
expect(out.content).toContain("$1::BIGINT");
expect(out.args.arg1).toBe("9007199254740993");
// Round-trip through JSON without throwing — the original bug.
expect(() => JSON.stringify(out.args)).not.toThrow();
});
test("BigInt zero / negative / large", () => {
const sql = dt();
expect(sql`SELECT ${BigInt(0)}`.args.arg1).toBe("0");
expect(sql`SELECT ${BigInt(-1)}`.args.arg1).toBe("-1");
expect(sql`SELECT ${BigInt("99999999999999999999")}`.args.arg1).toBe(
"99999999999999999999"
);
});
test("Date is auto-tagged ::TIMESTAMPTZ and ISO-stringified", () => {
// Pre-fix: typeof Date === "object" → ::JSON, then PG cast chain
// worked accidentally for `${date}::timestamptz`. Now: explicit
// ::TIMESTAMPTZ + Date.toISOString() so plain `${date}` against a
// timestamptz column doesn't need a cast.
const sql = dt();
const d = new Date("2024-01-15T10:30:00.000Z");
const out = sql`SELECT ${d} AS t`;
expect(out.content).toContain("$1::TIMESTAMPTZ");
expect(out.args.arg1).toBe("2024-01-15T10:30:00.000Z");
expect(() => JSON.stringify(out.args)).not.toThrow();
});
test("non-finite Number is stringified for the executor", () => {
// JSON.stringify(NaN) and JSON.stringify(Infinity) both produce `null`,
// which silently became NULL in the database. The executor accepts
// "NaN" / "Infinity" / "-Infinity" via `Value::String → FLOAT8`
// (`f64::from_str`), so we send the special values as strings.
const sql = dt();
expect(sql`SELECT ${NaN}`.args.arg1).toBe("NaN");
expect(sql`SELECT ${Infinity}`.args.arg1).toBe("Infinity");
expect(sql`SELECT ${-Infinity}`.args.arg1).toBe("-Infinity");
// Tag stays DOUBLE PRECISION (these are floats).
expect(sql`SELECT ${NaN}`.content).toContain("$1::DOUBLE PRECISION");
});
test("homogeneous arrays auto-tag with TYPE[]", () => {
const sql = dt();
expect(sql`SELECT ${[1, 2, 3]}`.content).toContain("$1::BIGINT[]");
expect(sql`SELECT ${[1.5, 2.5]}`.content).toContain(
"$1::DOUBLE PRECISION[]"
);
expect(sql`SELECT ${["a", "b"]}`.content).toContain("$1::TEXT[]");
expect(sql`SELECT ${[true, false]}`.content).toContain("$1::BOOLEAN[]");
});
test("non-homogeneous and empty arrays fall back to JSON", () => {
const sql = dt();
expect(sql`SELECT ${[1, "x"]}`.content).toContain("$1::JSON");
expect(sql`SELECT ${[]}`.content).toContain("$1::JSON");
expect(sql`SELECT ${[[1], [2]]}`.content).toContain("$1::JSON");
});
test("mixed-numeric array widens to DOUBLE PRECISION[]", () => {
const sql = dt();
expect(sql`SELECT ${[1, 2.5]}`.content).toContain(
"$1::DOUBLE PRECISION[]"
);
});
test("multiple args get distinct decls + numbered placeholders", () => {
const sql = dt();
const out = sql`INSERT INTO t VALUES (${1}, ${"x"}, ${[true, false]})`;
expect(out.content).toContain("-- $1 arg1");
expect(out.content).toContain("-- $2 arg2");
expect(out.content).toContain("-- $3 arg3");
expect(out.content).toContain("$1::BIGINT");
expect(out.content).toContain("$2::TEXT");
expect(out.content).toContain("$3::BOOLEAN[]");
expect(out.args).toMatchObject({
arg1: 1,
arg2: "x",
arg3: [true, false],
});
});
test("RawSql is inlined verbatim, doesn't consume an arg index", () => {
const sql = dt();
const col = new RawSql("name");
const out = sql`SELECT ${col} FROM t WHERE id = ${42}`;
// Only one decl, only one arg in args dict.
expect(out.content.match(/^-- \$\d+/gm)?.length).toBe(1);
expect(out.content).toContain("SELECT name FROM t WHERE id = $1::BIGINT");
expect(Object.keys(out.args).filter((k) => k.startsWith("arg")).length).toBe(
1
);
expect(out.args).toMatchObject({ arg1: 42 });
});
test("schema name is propagated as SET search_path preamble", () => {
const sql = dt("main");
const out = sql`SELECT 1`;
expect(out.args.database).toBe("datatable://main");
});
test("database extra arg is always present", () => {
const sql = dt("custom_db");
const out = sql`SELECT ${1}`;
expect(out.args.database).toBe("datatable://custom_db");
});
});
// =============================================================================
// datatable().query() — positional placeholders (the previously-broken path)
// =============================================================================
describe("datatable().query() — positional placeholders", () => {
test("emits typed declarations + SQL verbatim, no appended placeholders", () => {
const q = datatableQuery();
const out = q("SELECT $1, $2", 42, "hello");
expect(out.content).toContain("-- $1 arg1 (BIGINT)");
expect(out.content).toContain("-- $2 arg2 (TEXT)");
// Crucially: SQL must end with the user's SQL, NOT have placeholders
// appended after it (the pre-fix bug).
expect(out.content.endsWith("SELECT $1, $2")).toBe(true);
expect(out.args).toMatchObject({ arg1: 42, arg2: "hello" });
});
test("BigInt args are stringified", () => {
const q = datatableQuery();
const out = q("SELECT $1", BigInt("100"));
expect(out.args.arg1).toBe("100");
expect(out.content).toContain("-- $1 arg1 (BIGINT)");
});
test("array args auto-tag homogeneously in the decl block", () => {
const q = datatableQuery();
const out = q("SELECT $1, $2", [1, 2], ["a", "b"]);
expect(out.content).toContain("-- $1 arg1 (BIGINT[])");
expect(out.content).toContain("-- $2 arg2 (TEXT[])");
});
test("zero params → no decl block, just SQL + extras", () => {
const q = datatableQuery();
const out = q("SELECT 1");
expect(out.content).not.toContain("-- $");
expect(out.content).toContain("SELECT 1");
// database extra still injected.
expect(out.args.database).toBe("datatable://main");
});
test("ten params number contiguously", () => {
const q = datatableQuery();
const params = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10];
const out = q("SELECT " + params.map((_, i) => `$${i + 1}`).join(","), ...params);
for (let i = 1; i <= 10; i++) {
expect(out.content).toContain(`-- $${i} arg${i} (BIGINT)`);
expect(out.args[`arg${i}`]).toBe(i);
}
});
test(".query()'s decl format matches what the executor parses (regression)", () => {
// The PG parser's RE_ARG_PGSQL is:
// ^-- \$(\d+) (\w+)(?: \(([A-Za-z0-9_\[\]]+)\))?(?: ?\= ?(.+))? *$
// We assert our decl line matches that grammar so the executor doesn't
// fall back to "inferred default text".
const q = datatableQuery();
const out = q("SELECT $1", 42);
const declRe = /^-- \$\d+ \w+ \([A-Za-z0-9_\[\]]+\) *$/m;
expect(out.content).toMatch(declRe);
});
});
// =============================================================================
// ducklake() template tag — DuckDB declares types in the comment (different
// shape from datatable). Same auto-tag rules apply for inferSqlType.
// =============================================================================
describe("ducklake() — DuckDB shape", () => {
test("declarations carry the type", () => {
const sql = dl("main");
const out = sql`SELECT ${42}, ${"hello"}, ${true}`;
expect(out.content).toContain("-- $arg1 (BIGINT)");
expect(out.content).toContain("-- $arg2 (TEXT)");
expect(out.content).toContain("-- $arg3 (BOOLEAN)");
// Preamble attaches the ducklake.
expect(out.content).toContain("ATTACH 'ducklake://main' AS dl;USE dl;");
// Args are referenced via $argN syntax in the SQL body.
expect(out.content).toContain("$arg1");
});
test("BigInt + homogeneous arrays propagate to ducklake too", () => {
const sql = dl("main");
const out = sql`SELECT ${BigInt(9)}, ${[1, 2, 3]}, ${["a", "b"]}`;
expect(out.content).toContain("(BIGINT)");
expect(out.content).toContain("(BIGINT[])");
expect(out.content).toContain("(TEXT[])");
expect(out.args.arg1).toBe("9");
expect(out.args.arg2).toEqual([1, 2, 3]);
});
test("ducklake doesn't carry a database extra arg", () => {
const sql = dl();
const out = sql`SELECT 1`;
expect(out.args).not.toHaveProperty("database");
});
});
// =============================================================================
// Cross-cutting: the args dict must always be JSON-serialisable.
// =============================================================================
describe("args dict is JSON-serialisable for every supported value shape", () => {
test("BigInt, primitives, arrays, objects, raw — none throw", () => {
const sql = dt();
const out = sql`
SELECT ${BigInt(1)}, ${1}, ${1.5}, ${"x"}, ${true}, ${null},
${[1, 2]}, ${["a", "b"]}, ${[true, false]},
${{ k: 1 }}, ${[1, "x"]}
`;
const json = JSON.stringify(out.args);
expect(typeof json).toBe("string");
// BigInt got stringified, not thrown.
expect(json).toContain('"arg1":"1"');
});
});