mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-08-23 00:00:33 +00:00
all
This commit is contained in:
@@ -0,0 +1,219 @@
|
||||
|
||||
▐▛███▜▌ Claude Code v2.1.12
|
||||
▝▜█████▛▘ Opus 4.5 · Claude Max
|
||||
▘▘ ▝▝ ~/windmill/backend
|
||||
|
||||
❯ /clear
|
||||
⎿ (no content)
|
||||
|
||||
❯ If you look at how job assignment is done, it's done based on a tag which
|
||||
allow workers to do a super efficient pull. It works well and nice but I
|
||||
wonder if it would be possible to do a constraint assignment one instead
|
||||
where workers declare some property, a,b,c,d with some value, and job
|
||||
says that a job has constraints in terms of a,b,c,d with and or query. If
|
||||
it's impossible to make fast, maybe we could say that it's strict equal
|
||||
and always AND
|
||||
|
||||
● This is an interesting architectural question. Let me explore how the current
|
||||
tag-based job assignment works to understand the performance implications and
|
||||
feasibility of a constraint-based approach.
|
||||
|
||||
● Explore(Explore job assignment mechanism)
|
||||
⎿ Done (35 tool uses · 71.8k tokens · 1m 5s)
|
||||
|
||||
● Based on my exploration, here's how the current system works and my analysis
|
||||
of constraint-based assignment:
|
||||
|
||||
Current Tag System
|
||||
|
||||
The efficiency comes from:
|
||||
1. Single tag column on v2_job_queue
|
||||
2. B-tree index on (priority DESC, scheduled_for, tag) WHERE running = false
|
||||
3. Simple query: WHERE running = false AND tag IN ('tag1', 'tag2', ...) AND
|
||||
scheduled_for <= now()
|
||||
4. FOR UPDATE SKIP LOCKED for race-free claiming
|
||||
|
||||
The IN clause with a B-tree index is extremely efficient because PostgreSQL
|
||||
can do index-only scans.
|
||||
|
||||
Constraint-Based Assignment Options
|
||||
|
||||
Option 1: JSONB containment (AND/OR support)
|
||||
|
||||
-- Workers store properties
|
||||
ALTER TABLE worker_ping ADD COLUMN properties JSONB;
|
||||
-- Jobs store constraints
|
||||
ALTER TABLE v2_job_queue ADD COLUMN constraints JSONB;
|
||||
|
||||
-- Pull query becomes
|
||||
SELECT id FROM v2_job_queue
|
||||
WHERE running = false
|
||||
AND (constraints IS NULL OR worker_properties @> constraints)
|
||||
ORDER BY priority DESC NULLS LAST, scheduled_for
|
||||
FOR UPDATE SKIP LOCKED
|
||||
|
||||
Problem: This inverts the query direction. Currently workers pull jobs
|
||||
matching their tags. With JSONB, you'd need to check if worker properties
|
||||
satisfy job constraints, but the worker properties aren't in the job table.
|
||||
You'd need a JOIN or subquery, killing performance.
|
||||
|
||||
Option 2: Strict AND with equality - Constraint Hash
|
||||
|
||||
If constraints are always AND with strict equality, you can compute a
|
||||
canonical hash:
|
||||
|
||||
// Job with constraints {a: "gpu", b: "high-mem"}
|
||||
// Gets constraint_hash = hash("a=gpu&b=high-mem")
|
||||
|
||||
// Worker with properties {a: "gpu", b: "high-mem", c: "fast-disk"}
|
||||
// Pre-computes ALL hashes it can satisfy:
|
||||
// - hash("") (no constraints)
|
||||
// - hash("a=gpu")
|
||||
// - hash("b=high-mem")
|
||||
// - hash("c=fast-disk")
|
||||
// - hash("a=gpu&b=high-mem")
|
||||
// - hash("a=gpu&c=fast-disk")
|
||||
// - ... (all 2^n combinations)
|
||||
|
||||
Pull query stays efficient:
|
||||
WHERE constraint_hash IN ('hash1', 'hash2', ...)
|
||||
|
||||
Downside: 2^n hashes per worker. With 4 properties = 16 hashes (manageable).
|
||||
With 8 = 256 (getting large).
|
||||
|
||||
Option 3: Fixed constraint columns (pragmatic)
|
||||
|
||||
ALTER TABLE v2_job_queue ADD COLUMN constraint_gpu VARCHAR;
|
||||
ALTER TABLE v2_job_queue ADD COLUMN constraint_memory VARCHAR;
|
||||
ALTER TABLE v2_job_queue ADD COLUMN constraint_region VARCHAR;
|
||||
|
||||
CREATE INDEX queue_constraints ON v2_job_queue
|
||||
(running, constraint_gpu, constraint_memory, constraint_region, priority
|
||||
DESC, scheduled_for)
|
||||
WHERE running = false;
|
||||
|
||||
Workers generate query:
|
||||
WHERE running = false
|
||||
AND (constraint_gpu IS NULL OR constraint_gpu = 'nvidia-a100')
|
||||
AND (constraint_memory IS NULL OR constraint_memory IN ('high', 'medium'))
|
||||
...
|
||||
|
||||
Downside: The IS NULL OR pattern doesn't index-scan as cleanly. PostgreSQL
|
||||
might fall back to bitmap scans.
|
||||
|
||||
Option 4: Keep tags, but make them composite
|
||||
|
||||
Instead of changing the architecture, allow composite tags:
|
||||
tag = "gpu:a100,mem:high,region:us-east"
|
||||
|
||||
Workers declare which composite tags they handle. Jobs specify their required
|
||||
composite tag.
|
||||
|
||||
Downside: Loses flexibility - can't match partial constraints.
|
||||
|
||||
My Recommendation
|
||||
|
||||
For strict AND with equality, the constraint hash approach (Option 2) is the
|
||||
cleanest because it preserves the current query structure. The 2^n blowup is
|
||||
acceptable if you limit to ~5-6 constraint dimensions.
|
||||
|
||||
Would you like me to prototype what this would look like in the codebase? I
|
||||
could show:
|
||||
1. Schema changes
|
||||
2. How workers would compute their satisfiable constraint hashes
|
||||
3. How jobs would compute their constraint hash at creation time
|
||||
|
||||
✻ Cooked for 2m 5s
|
||||
|
||||
❯ couldn't hash be a tag in the end? What if I say that at most 3
|
||||
properties can be set globally ?
|
||||
|
||||
● Exactly right. With 3 properties max, workers compute at most 2³ = 8
|
||||
constraint tags. The existing system stays unchanged:
|
||||
|
||||
Example:
|
||||
|
||||
Global properties: gpu, mem, region
|
||||
|
||||
Worker declares: {gpu: "a100", mem: "high"} (no region)
|
||||
|
||||
Worker generates these tags:
|
||||
c: # satisfies jobs with no constraints
|
||||
c:gpu=a100
|
||||
c:mem=high
|
||||
c:gpu=a100,mem=high
|
||||
|
||||
Job with constraints {gpu: "a100", mem: "high"} gets tag:
|
||||
c:gpu=a100,mem=high
|
||||
|
||||
The existing tag IN (...) query works unchanged. No schema changes, no index
|
||||
changes.
|
||||
|
||||
Key implementation details:
|
||||
|
||||
1. Canonical ordering - always sort keys alphabetically so {mem: "high", gpu:
|
||||
"a100"} produces the same tag as {gpu: "a100", mem: "high"}
|
||||
2. Workers without a property can't satisfy jobs requiring it - they simply
|
||||
don't generate tags containing that property
|
||||
3. Jobs without constraints get tag c: (or could coexist with regular tags
|
||||
like deno, python3)
|
||||
4. Coexistence - constraint tags could be a separate dimension from language
|
||||
tags, or you could combine them: deno,c:gpu=a100,mem=high
|
||||
|
||||
Want me to sketch out what the tag generation logic would look like in Rust?
|
||||
|
||||
❯ create a feature branch, put your notes on it, then push that branch
|
||||
|
||||
● Bash(cd /home/rfiszel/windmill && git checkout -b
|
||||
feature/constraint-based-job-assignment)
|
||||
⎿ Switched to a new branch 'feature/constraint-based-job-assignment'
|
||||
⎿ Shell cwd was reset to /home/rfiszel/windmill/backend
|
||||
|
||||
● Write(docs/design/constraint-based-job-assignment.md)
|
||||
⎿ Wrote 265 lines to docs/design/constraint-based-job-assignment.md
|
||||
1 # Constraint-Based Job Assignment
|
||||
2 ## Overview
|
||||
3 This document explores adding constraint-based job assignment to
|
||||
Windmill, allowing workers to declare properties and jobs to spec
|
||||
ify constraints that must be satisfied.
|
||||
4 ## Current System
|
||||
5 Jobs are assigned based on a single `tag` column in `v2_job_queue
|
||||
`. Workers declare which tags they handle, and pull jobs using:
|
||||
… +255 lines (ctrl+o to expand)
|
||||
|
||||
● Bash(cd /home/rfiszel/windmill && git add
|
||||
backend/docs/design/constraint-based-job-assignment.md && git commit -m
|
||||
"docs: add design document for constraint-based job…)
|
||||
⎿ [feature/constraint-based-job-assignment 8c34c65d8e] docs: add design d
|
||||
ocument for constraint-based job assignment
|
||||
1 file changed, 264 insertions(+)
|
||||
… +2 lines (ctrl+o to expand)
|
||||
|
||||
⎿ Shell cwd was reset to /home/rfiszel/windmill/backend
|
||||
|
||||
● Bash(cd /home/rfiszel/windmill && git push -u origin
|
||||
feature/constraint-based-job-assignment)
|
||||
⎿ branch 'feature/constraint-based-job-assignment' set up to track 'origi
|
||||
n/feature/constraint-based-job-assignment'.
|
||||
remote:
|
||||
… +14 lines (ctrl+o to expand)
|
||||
|
||||
⎿ Shell cwd was reset to /home/rfiszel/windmill/backend
|
||||
|
||||
● Done. Branch feature/constraint-based-job-assignment pushed with the design
|
||||
document at backend/docs/design/constraint-based-job-assignment.md.
|
||||
|
||||
The document covers:
|
||||
- Current tag system analysis
|
||||
- Constraint tag format (c:gpu=a100,mem=high)
|
||||
- Worker tag generation (all 2^n combinations)
|
||||
- Schema changes needed
|
||||
- Rust implementation sketch for tag generation
|
||||
- Query modification
|
||||
- Performance analysis
|
||||
|
||||
PR link: https://github.com/windmill-labs/windmill/pull/new/feature/constraint
|
||||
-based-job-assignment
|
||||
|
||||
✻ Crunched for 54s
|
||||
|
||||
Reference in New Issue
Block a user