Files
windmill/backend/windmill-api-schedule/src
hugocasaandClaude Opus 5.5 97fa55b719 feat: detect and alert when a schedule skips occurrences (#10917)
* docs: plan for detecting skipped schedule occurrences

Design plan only, no implementation. Records the scheduler's re-anchoring
behaviour, the measurements behind it, and the three-piece design that came
out of reviewing the alternatives.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* docs: state the user-facing outcome in the schedule plan

The plan described the mechanism but never what a user would see, which made it
hard to judge what the work is worth.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* docs: state which cause the schedule plan catches, and correct its scope

Records which of the two causes each piece covers, and corrects the overrun
scope: a script schedule carrying retry or dynamic_skip is pushed as a
SingleStepFlow, so it re-arms at step 0 entry and its occurrences overlap like
a flow's. Resolves the no_flow_overlap question, splits the read-time work into
bounded detection and editor-only counting behind measured croner costs, and
fixes the delivery order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: count the occurrences a schedule skipped

A schedule that overruns its interval, or waits for a worker, silently loses
the occurrences in between: the scheduler keeps one queued occurrence and
re-anchors on the clock, so nothing records that a run was due and never
happened.

Recovers the sequence from rows that already exist rather than writing per
occurrence. `push_scheduled_job` anchors on `now_from_db` inside the
transaction that inserts the job, and `v2_job.created_at` defaults to that
same transaction timestamp, so `scheduled_for = find_next(created_at)` holds
exactly and the whole occurrence history is derivable.

The schedules list reports how many of the recent runs were followed by a lost
occurrence, and a new occurrences endpoint carries the per-run wait and
duration behind it. Detection is one `find_next` per gap, which stays bounded
on a full page; counting walks the gap and runs only for a single schedule.

The one write is `occurrence_baseline_at`, advanced at create, edit,
re-enable and re-arm. Gaps older than it span a pause, a cron change, a
re-enable or a reconciler re-arm, none of which mean runs were lost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: show the wait and run time behind a schedule's skipped occurrences

The list badge says a schedule is losing runs; this says which of the two
causes did it. A large wait means not enough workers, a long run means the job
outgrew its interval, and the pair is what tells them apart. Sits under the
existing upcoming-events panel, so due and overdue read together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: flag a schedule that is running late right now

Reconstruction is retrospective: a gap only appears once the next occurrence
has a row, which needs the current one to finish. A schedule wedged mid-run
shows nothing until it moves, which is the case an operator most wants to see.

An occurrence still in flight past the time its own successor was due will
cost that successor, so `now > find_next(scheduled_for)` is the signal, needing
no threshold and self-calibrating across a daily and a per-minute schedule. It
applies only where occurrences serialize; an overlapping schedule starts its
successor on time and would flag constantly while healthy.

The queue is read in one aggregating pass keyed on (trigger, runnable_path)
rather than a subquery per schedule, and an overlapping schedule holds more
than one root row, hence the aggregate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: run the schedule overrun alert from the monitor pass

Wires `schedule_overrun_alerts` in next to `jobs_waiting_alerts`, every 30
iterations (~5 min). Its Enterprise implementation lives in
windmill-labs/windmill-ee-private#772; only the wiring and the OSS stub are
here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* chore: refresh the sqlx offline cache

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: record and alert when a schedule skips occurrences

push_scheduled_job compares each chained occurrence with the slot after
the previous one. A gap is written to schedule.skipped_occurrences off the
push transaction, alerts once when a clean schedule starts skipping, and
recovers on the next clean chain. The schedules list shows a badge.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: alert only on a streak of skipping runs, keep a recent skip visible

The skip state now describes the current streak and is written in the push
transaction, so it commits or rolls back with the push. The alert fires
once when 3 runs in a row skipped, and the list keeps a muted badge for 7
days after the latest skip. Editing or toggling a schedule resets it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* fix: name the missed-occurrence state after what it counts, alert only once committed

Renames the columns to late_run_streak, missed_occurrences and
last_missed_at, keeps the missed count after a streak ends so the muted
badge can show it, and rewords both badges. The alert task now reads the
streak FOR SHARE, which waits for the push transaction, so a push that
rolls back and retries alerts once. A failed slot count leaves the streak
untouched, and a schedule deleted mid-push no longer fails it. Adds an
integration test for the streak and its reset.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* fix: recover the late run alert, store the missed slot, name it missed throughout

The alert now recovers (and so acknowledges itself) when a streak that
alerted ends on a run on time, under the schedule:{path} resource used by
the other trigger alerts. last_missed_at records the last missed cron slot
rather than when the late run chained, and the counting helpers say
missed, since skipped already names occurrences queued and not run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* fix: scope the late run alert to its workspace, acknowledge it on edit, toggle and delete

Recovery acknowledges alerts by resource alone, so the resource now
carries the workspace. Editing, toggling or deleting a schedule clears
its streak and a disabled or deleted one never chains a run on time, so
those handlers acknowledge its open alert after committing. Past the
1000-slot cap, last_missed_at falls back to the detection time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* refactor: raise the late run alert like the other critical alerts

Drops the recovery, the workspace-scoped resource and the acknowledgement
on edit, toggle and delete: the alert now fires once per streak with no
resource and is acknowledged from the alerts feed, as the trigger and job
failure alerts are. The FOR SHARE read stays, so a push that rolls back
across the flow path's retries still alerts once. Notes in openapi that
past 1000 misses in one late run the count is a floor and the time
approximate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-29 12:40:25 +02:00
..