mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-10-03 08:02:19 +00:00
* docs: plan for detecting skipped schedule occurrences Design plan only, no implementation. Records the scheduler's re-anchoring behaviour, the measurements behind it, and the three-piece design that came out of reviewing the alternatives. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * docs: state the user-facing outcome in the schedule plan The plan described the mechanism but never what a user would see, which made it hard to judge what the work is worth. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * docs: state which cause the schedule plan catches, and correct its scope Records which of the two causes each piece covers, and corrects the overrun scope: a script schedule carrying retry or dynamic_skip is pushed as a SingleStepFlow, so it re-arms at step 0 entry and its occurrences overlap like a flow's. Resolves the no_flow_overlap question, splits the read-time work into bounded detection and editor-only counting behind measured croner costs, and fixes the delivery order. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: count the occurrences a schedule skipped A schedule that overruns its interval, or waits for a worker, silently loses the occurrences in between: the scheduler keeps one queued occurrence and re-anchors on the clock, so nothing records that a run was due and never happened. Recovers the sequence from rows that already exist rather than writing per occurrence. `push_scheduled_job` anchors on `now_from_db` inside the transaction that inserts the job, and `v2_job.created_at` defaults to that same transaction timestamp, so `scheduled_for = find_next(created_at)` holds exactly and the whole occurrence history is derivable. The schedules list reports how many of the recent runs were followed by a lost occurrence, and a new occurrences endpoint carries the per-run wait and duration behind it. Detection is one `find_next` per gap, which stays bounded on a full page; counting walks the gap and runs only for a single schedule. The one write is `occurrence_baseline_at`, advanced at create, edit, re-enable and re-arm. Gaps older than it span a pause, a cron change, a re-enable or a reconciler re-arm, none of which mean runs were lost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: show the wait and run time behind a schedule's skipped occurrences The list badge says a schedule is losing runs; this says which of the two causes did it. A large wait means not enough workers, a long run means the job outgrew its interval, and the pair is what tells them apart. Sits under the existing upcoming-events panel, so due and overdue read together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: flag a schedule that is running late right now Reconstruction is retrospective: a gap only appears once the next occurrence has a row, which needs the current one to finish. A schedule wedged mid-run shows nothing until it moves, which is the case an operator most wants to see. An occurrence still in flight past the time its own successor was due will cost that successor, so `now > find_next(scheduled_for)` is the signal, needing no threshold and self-calibrating across a daily and a per-minute schedule. It applies only where occurrences serialize; an overlapping schedule starts its successor on time and would flag constantly while healthy. The queue is read in one aggregating pass keyed on (trigger, runnable_path) rather than a subquery per schedule, and an overlapping schedule holds more than one root row, hence the aggregate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: run the schedule overrun alert from the monitor pass Wires `schedule_overrun_alerts` in next to `jobs_waiting_alerts`, every 30 iterations (~5 min). Its Enterprise implementation lives in windmill-labs/windmill-ee-private#772; only the wiring and the OSS stub are here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * chore: refresh the sqlx offline cache Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: record and alert when a schedule skips occurrences push_scheduled_job compares each chained occurrence with the slot after the previous one. A gap is written to schedule.skipped_occurrences off the push transaction, alerts once when a clean schedule starts skipping, and recovers on the next clean chain. The schedules list shows a badge. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: alert only on a streak of skipping runs, keep a recent skip visible The skip state now describes the current streak and is written in the push transaction, so it commits or rolls back with the push. The alert fires once when 3 runs in a row skipped, and the list keeps a muted badge for 7 days after the latest skip. Editing or toggling a schedule resets it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * fix: name the missed-occurrence state after what it counts, alert only once committed Renames the columns to late_run_streak, missed_occurrences and last_missed_at, keeps the missed count after a streak ends so the muted badge can show it, and rewords both badges. The alert task now reads the streak FOR SHARE, which waits for the push transaction, so a push that rolls back and retries alerts once. A failed slot count leaves the streak untouched, and a schedule deleted mid-push no longer fails it. Adds an integration test for the streak and its reset. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * fix: recover the late run alert, store the missed slot, name it missed throughout The alert now recovers (and so acknowledges itself) when a streak that alerted ends on a run on time, under the schedule:{path} resource used by the other trigger alerts. last_missed_at records the last missed cron slot rather than when the late run chained, and the counting helpers say missed, since skipped already names occurrences queued and not run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * fix: scope the late run alert to its workspace, acknowledge it on edit, toggle and delete Recovery acknowledges alerts by resource alone, so the resource now carries the workspace. Editing, toggling or deleting a schedule clears its streak and a disabled or deleted one never chains a run on time, so those handlers acknowledge its open alert after committing. Past the 1000-slot cap, last_missed_at falls back to the detection time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * refactor: raise the late run alert like the other critical alerts Drops the recovery, the workspace-scoped resource and the acknowledgement on edit, toggle and delete: the alert now fires once per streak with no resource and is acknowledged from the alerts feed, as the trigger and job failure alerts are. The FOR SHARE read stays, so a push that rolls back across the flow path's retries still alerts once. Notes in openapi that past 1000 misses in one late run the count is a floor and the time approximate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>