WorkerHealthAlert lives in the admin layout and renders on every admin
page when any worker needs attention. Polls the workers list every 30s
in the background. Hidden on the workers page itself (which has its own
richer banner with filter chips).
Surfaces three categories:
- errored: install_state == "error"
- offline: installed + no heartbeat for 5+ minutes
- in progress: pending / provisioning / uninstalling
Tone goes red when there are errored or offline workers, amber otherwise.
A "review →" link deep-jumps to /app/admin/workers.
New /app/admin/audit page browses admin_audit_log with:
- Filter by action, target_type, target_id, admin_user_id, date range
- Action + target_type dropdowns are populated from a baseline set
AND from whatever appears in the current result (so new actions
surface automatically without code changes)
- Auto-refresh toggle (5s) for live tailing during operations
- Cursor-based pagination with Prev/Next + page counter
- Expandable rows show the details JSON + user-agent
- Color-coded actions (red for destructive, green for create/install,
amber for system auto-actions)
- Renders system actions (admin_user_id = uuid.Nil) as "system"
- Renders the admin's name + email when joined data is present
When the dead-worker job reassigns email accounts from a worker whose
heartbeat expired, write a row into admin_audit_log so the dashboard's
audit viewer surfaces these system actions alongside admin-driven ones.
admin_user_id is uuid.Nil (the platform identity), so admins searching
the log can distinguish "system did this" from "an admin did this" by
filtering on that ID. Details include the replacement worker, account
count, and reason.
JobsService gets an optional AdminRepo dep. Nil disables logging — keeps
the contract loose for any other call site that doesn't have one.
The audit calls added in the last commit went to AuditService.LogAction,
which writes to the general user-facing audit log (Cassandra). The admin
audit-log viewer at /admin/audit-logs queries the admin_audit_log table
in Postgres, so worker / credentials / release actions never showed up.
Add a public AdminService.LogAdminAction that wraps the existing private
logAction (writes to admin_audit_log with the same shape as ban_user /
update_worker / etc.). Repoint h.audit() at it.
Actions now visible in the audit viewer:
test, install, restart, upgrade, uninstall, rotate_keys, apply,
assign, system_update, reboot, check_releases (plus the existing
create/update/delete across workers, AWS creds, and profiles).
The seeder is one of the few things every developer runs on every new
checkout, but it had zero tests. With three migrations added in the last
few days and the rich-fixture path now creating 30+ rows, the chance of
silently breaking a schema migration without noticing was non-trivial.
cmd/seed/main_test.go connects to SEED_TEST_DB (skips otherwise — keeps
unit tests in CI fast and prevents accidentally clobbering a dev
database), runs migrations, wipes only the fixture rows, then:
1. Runs seedBaseline twice, verifies row count stays at 1.
2. Runs seedRich, asserts 9 different row counts match expectations
(users, orgs, workers, accounts, campaign, sequences, contacts,
unsubscribed contacts, campaign leads).
3. Re-runs seedRich, asserts every count is unchanged — the most
important guarantee the seeder makes.
4. Verifies warmup pool membership: 2 free, 4 premium, with the
correct accounts in each.
Plain testing package, table-driven, matches the existing style in
internal/app/warmup/service_test.go.
`make test-seed` brings up the docker-compose Postgres and runs the
suite against it.
Workers list now computes a health summary client-side from the existing
queries and shows a yellow banner with a count when any worker needs
attention (errored, offline, or mid-install). Filter chips below the
banner narrow the table by problem category so the admin can drill in
without opening each worker.
Categories surfaced:
- errored: install_state = "error"
- offline: installed, but no heartbeat in 5+ minutes
- in progress: pending / provisioning / uninstalling
- stale config: profile.updated_at > worker.config_applied_at
- update available: profile.resolved_image_tag != worker.image_version
Live polling stays at 15s, so a worker going offline shows up within
~75s of the last missed heartbeat. The list also fetches profiles so the
stale-config and update-available chips can compute without extra
round-trips.
The new admin endpoints for SSH worker management, AWS credentials,
worker profiles, releases, and system operations bypassed the existing
audit-logging pattern. Now every mutating action records a row in the
audit log with adminID, IP, user-agent, and operation-specific metadata
(never secret values — for credential updates we record which fields
were rotated, not what they became).
New action constants:
test, install, restart, upgrade, uninstall, rotate_keys, apply,
assign, system_update, reboot, check_releases
New entity types:
worker, aws_credentials, worker_profile, release
Read-only endpoints (list/get/status/logs) intentionally not audited —
they don't change state and would flood the log.
Each handler calls a small h.audit(c, action, entity, &id, metadata)
helper that pulls adminID from the admin middleware and fires the
existing AuditService.LogAction (fire-and-forget, never blocks the
response).
Workers heartbeat into Redis every 90s as RFC3339 timestamp values with a
3-min TTL. The dashboard surfaces liveness based on workers.last_seen_at,
but until now nothing populated that column — the "Live" badge was always
red.
New 60s job in the consumer reads each active worker's Redis heartbeat
value, parses the timestamp, and writes it to workers.last_seen_at. Runs
on its own interval (separate from the 5-min dead-worker detection job,
which does heavier reassignment work) so the UI sees fresh data within a
minute.
Old docs described a k8s/ArgoCD/Terraform deployment that no longer
exists, with ASCII-art system diagrams that hadn't aged well. Rewritten
to match how the project actually ships:
- README: control plane (Railway) + execution plane (per-VPS workers)
split, dashboard-driven worker management, credentials/profiles,
auto-update from GitHub releases, OS package updates, self-hosting
knobs. Removed all ASCII art.
- resources/architecture.md: control vs execution plane, encryption
model (worker SSH keys + platform secrets under the same KMS-envelope
cipher as user secrets), worker identity from public IPv4, credentials
model, push-driven release flow, anti-abuse layers, source anchors.
- resources/deployment-guide.md: end-to-end from "provision a VPS" to
"auto-update on release". No more k8s, ArgoCD, kubectl, or Terraform.
Step-by-step backend env, webhook setup, worker add flow, day-2 ops,
rollback per plane.
- resources/local-development.md: the five make targets (dev / sim /
seed / tools / reset), what each profile runs, LocalStack bootstrap,
rich seed contents, native dev against containerized infra, the
offset-port URL table.
- resources/cicd.md: the two-plane build/release flow, image tag scheme
({sha} / dev / vX.Y.Z / vX.Y / vX / prod), webhook setup, release
process, security notes around HMAC and least-privilege worker AWS
keys.
- deploy/README.md: tight version of the same.
Workers are no longer curl|sh-only. Admins add and manage them from the
dashboard over SSH, with all runtime config (Kafka, Schema Registry,
Redis, AWS keys) stored encrypted via the existing KMS-envelope cipher
service.
Worker lifecycle:
1. Admin POSTs host/port/user. Backend generates an ed25519 keypair,
encrypts the private key under uuid.Nil (platform identity), and
stores the row in 'pending' state.
2. Admin pastes the returned public key into the VPS's authorized_keys.
3. Test connection — runs `true` over SSH, pins the host SHA256
fingerprint on first success (TOFU).
4. Install — backend scp's install-worker.sh + a per-worker env file
and runs it. State moves pending → provisioning → installed.
5. From then on: restart, update image, apply config, uninstall,
rotate keys, tail logs, live status, OS package update, reboot —
all dashboard buttons backed by SSH operations.
Credentials are reusable entities:
- aws_credentials: named keypair, secret encrypted at rest
- worker_profiles: bundles Kafka + Schema Registry + Redis + image +
release channel, references one AWS credentials row
- workers.profile_id links a worker to a profile; many workers can
share one profile
Saving a profile doesn't restart anything. The dashboard compares
profile.updated_at to each worker's config_applied_at and shows a
"stale config" badge; Apply rewrites /etc/warmbly/worker.env over SSH
and restarts the unit.
Auto-update on GitHub release:
- profile.release_channel ∈ {pinned, stable, dev}
- profile.auto_update toggles automatic rollout
- Trigger model is push, not poll: one check on backend boot, then
the /webhooks/github/releases endpoint (HMAC-validated with
RELEASES_WEBHOOK_SECRET) on every release event. Manual "Check now"
button as fallback.
- When a new tag resolves, the orchestrator SSHes into each assigned
worker, runs install-worker.sh --update --image <new>, which now
rewrites the systemd unit (not just `docker pull`) so the image
actually changes. workers.image_version captures the running tag
for the UI's "v1.2.3 → v1.2.4" diff.
Self-hostable: every release knob is env-driven —
RELEASES_GITHUB_REPO, RELEASES_WORKER_IMAGE_REPO,
RELEASES_WEBHOOK_SECRET, RELEASES_GITHUB_TOKEN, RELEASES_ENABLED. Set
RELEASES_ENABLED=false to disable the feature entirely.
OS-level updates and reboot are also exposed: detect apt / dnf / yum /
pacman / apk, run the right upgrade noninteractively, return the full
output and a reboot-required flag. Reboots are never automatic.
Migrations:
000028_worker_ssh — ssh fields, install_state enum, last_seen,
host fingerprint
000029_worker_credentials — aws_credentials + worker_profiles +
workers.profile_id + workers.config_applied_at
000030_worker_releases — release_channel enum, auto_update,
resolved_image_tag, workers.image_version
Endpoints added:
POST /admin/workers (create + keypair)
GET /admin/workers/managed
GET /admin/workers/:id/managed
POST /admin/workers/:id/{test,install,restart,upgrade,uninstall,rotate-keys,apply,system-update,reboot}
PUT /admin/workers/:id/profile
GET /admin/workers/:id/{live-status,logs}
DELETE /admin/workers/:id
GET /admin/aws-credentials CRUD
GET /admin/worker-profiles CRUD + /workers + /apply + /release
GET /admin/releases/state
POST /admin/releases/check
POST /webhooks/github/releases public, HMAC-validated
Admin UI:
/app/admin/workers list with status + version columns
/app/admin/workers/new add form with profile dropdown
/app/admin/workers/:id detail with all actions + logs + system update
/app/admin/credentials tabs: AWS credentials + worker profiles,
Releases panel, channel selector +
auto-update toggle in profile form
Hoist the dev/sim stack to a single docker-compose.yml at the repo root.
Adds profiles (default / sim / seed / tools) so you can opt into heavier
setups, and bundles dependencies that were previously missing:
- LocalStack (KMS + DynamoDB + S3) with a localstack-init one-shot that
idempotently creates alias/master-key-dev, the UserEncryptedKeys and
EmailMessageData tables, and the main S3 bucket. Backend and workers
wait on it via service_completed_successfully.
- stripe-mock for billing flows
- kafka-ui under the tools profile
Three workers with deterministic UUIDv5 hostnames (shared / premium /
dedicated) so assignment, rebalancing, and per-pool routing all have
real targets to exercise.
Richer seed (cmd/seed/main.go) loads 3 orgs across tiers, 6 mailboxes
joined to free/premium warmup pools, a Beta campaign with a 2-step
sequence, and 10 contacts (2 unsubscribed) so suppression behaviour is
visible in the UI. Idempotent — safe to re-run.
Makefile targets:
make dev — infra + app + one worker
make sim — adds premium + dedicated workers
make seed — rich fixtures
make tools — kafka-ui at :18090
make reset — nuke volumes
scripts/install-worker.sh is a single bash script any Debian/Ubuntu/RHEL/
Fedora/Arch/Alpine VPS can curl|sh to add a worker to the fleet.
Identity is bound to the VPS's public IPv4 via UUIDv5 (URL namespace):
same IP → same worker (reputation persists across reinstalls)
new IP → new worker (fresh identity, no inherited reputation)
The installer detects the public IP via api.ipify.org / ifconfig.me /
checkip.amazonaws.com, derives the deterministic UUID, installs Docker if
missing, writes /etc/warmbly/worker.env (0600) and /etc/warmbly/worker.id,
installs a systemd unit that runs the worker container with --hostname
<uuid>, and starts the service.
Supports --install/--update/--uninstall/--purge/--status, --env-file for
non-interactive config, --ip override, --image override, and a full set
of per-credential flags.
Worker reads its UUID from os.Hostname() at startup, so the systemd
hostname value becomes the worker identity — no separate registration
step needed.
Workers need IP diversity, but k8s nodes typically NAT all pods through a
small set of egress IPs — defeating the point of a DaemonSet for cold mail.
Plus, the control plane is moving to Railway and workers will be managed
per-VPS, so the kustomize tree no longer reflects how anything actually
ships.
- cipher: use the already-available ctx parameter for DynamoDB Put
- goog: use context.Background for OAuth token refresh callback since
it runs asynchronously outside any request lifecycle
The Tries counter on login and registration sessions was checked but
never incremented, making the brute-force protection dead code. An
attacker could retry verification codes indefinitely within the session
TTL. Now each failed attempt increments and persists the counter.