mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-08-18 16:02:10 +00:00
feat: sandboxed daemonless container runtime via '# sandbox <image>' (#9453)
* feat: add sandboxed docker v2 runtime via '# docker <image>' Run a container image as a subprogram of the job's own nsjail sandbox: extract the image rootfs with podman (rootless) and run it chrooted inside the job's nsjail, so the container inherits the job's confinement and is safe under nsjail / for untrusted code. Selected by '# docker <image>'; a bare '# docker' keeps the v1 (dind) path untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: default to daemonless docker (drop dind from compose, allow docker on cloud) docker-compose no longer ships the dind sidecar (v2 is daemonless: podman + nsjail in the worker); removed the dind service, DOCKER_HOST env, depends_on and volume. Removed the language-picker guard that blocked Docker scripts on the multi-tenant platform, now that v2 makes docker safe to run sandboxed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: select sandboxed container via # sandbox <image>; add pull policy + size guards - Surface moved from '# docker <image>' to '# sandbox <image>' (groups under the sandbox annotation; '# docker' stays v1-only, '# sandbox' stays nsjail-bash). - SANDBOX_IMAGE_PULL_POLICY (default 'newer') so moving tags don't go stale. - SANDBOX_IMAGE_MAX_SIZE_MB rejects oversized images before extraction. - SANDBOX_IMAGE_CACHE_MAX_MB best-effort LRU eviction of podman's image store. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(sandbox): support # volume, honor nsjail tmp instance settings, v2 docker template - Thread shared_mount into the sandbox container nsjail config so '# volume' mounts (and the same-worker /tmp/shared folder) apply inside the container. - Use resolve_nsjail_tmp_mount_block for the container's /tmp so it honors the same nsjail_tmp_backing / nsjail_tmpfs_size_mb instance settings as other nsjail jobs. - docker-compose comment + the editor's Docker template now use '# sandbox <image>'. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(sandbox): make image size/cache/pull-policy UI instance settings Convert SANDBOX_IMAGE_* from worker env vars to DB-backed instance settings (sandbox_image_max_size_mb, sandbox_image_cache_max_mb, sandbox_image_pull_policy), hot-reloaded via the same mechanism as nsjail_tmpfs_size_mb and configurable in #superadmin-settings. No worker restart needed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(sandbox): windmill-managed registry — default registry + private auth Two new instance settings: - sandbox_image_default_registry: prepended to unqualified image refs (alpine -> <registry>/alpine); fully-qualified refs untouched. - sandbox_registry_auth: docker/podman auth.json blob written to a per-job authfile (0600, removed with the job) and passed to podman --authfile for private registries. Both hot-reloaded and configurable in #superadmin-settings. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sandbox): protobuf-safe proto_str escaper, atomic 0600 authfile, registry tests Addresses local-review P2s: proto_str now emits valid protobuf octal escapes for control/non-ASCII bytes (not Rust \u{..} that nsjail would reject); the registry authfile is created 0600 atomically (no world-readable window); add a registry_qualified table test + a non-ASCII proto_str case. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sandbox): P0 — deliver image env via nsjail envar:, never the launcher process env CI review (P0): the image's OCI Env (attacker-controlled keys+values) was applied to the nsjail launcher process via .envs(), so a hostile image could set LD_PRELOAD/ LD_LIBRARY_PATH/LD_AUDIT on nsjail itself and execute code as the worker outside the jail. Now the image env is rendered as proto-escaped 'envar:' directives (child-only) and nsjail's process env carries only windmill-trusted keys (reserved vars + proxy). Also: warn instead of silently bypassing the size guard on inspect failure; reset the eviction guard via a Drop guard (no stuck flag on panic/early-return). +render_envars test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sandbox): P0 symlink-write escape via rootfs script; P1 redact registry-auth logging CI review: - P0 (Codex): the body was written into the image-controlled rootfs as .windmill_docker_main.sh via write_file (follows symlinks) — a hostile image could plant that path as a symlink to a host file and capture the worker's write before nsjail starts. Now the body is passed straight to 'sh -c <body> sh <args>'; no file is written into the rootfs at all. - P1 (Codex): sandbox_registry_auth flowed through the generic setting loader which logs the value (raw auth.json credentials). Replaced with a secret-aware reload that loads directly and logs only a redacted 'configured=' message. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sandbox): redact sandbox_registry_auth in instance-settings write log too The settings API also logs 'Set global setting <key> to <value>' via format_setting_value; add sandbox_registry_auth to SENSITIVE_SETTINGS so the credential is redacted there as well as on reload. * fix(sandbox): don't silently disable cache eviction on podman images parse error Re-review (cubic/Claude P2): serde_json::from_slice(...).unwrap_or_default() meant any parse hiccup (e.g. podman omitting Size/Created via omitempty for a zero value, or schema drift) silently degraded to an empty Vec and disabled eviction with no log. Now Size/Created are #[serde(default)] (a missing omitempty key -> 0, not a whole-array parse failure) and a real parse error warns + breaks instead of being swallowed. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
+26
-7
@@ -57,11 +57,14 @@ use windmill_common::{
|
||||
PREVIEW_TAGS_OVERRIDE_SETTING, REQUEST_SIZE_LIMIT_SETTING,
|
||||
REQUIRE_PREEXISTING_USER_FOR_OAUTH_SETTING, RESTART_COORDINATION_SETTING,
|
||||
RETENTION_PERIOD_SECS_SETTING, RUBY_REPOS_SETTING, SAML_METADATA_SETTING,
|
||||
SCIM_TOKEN_SETTING, SMTP_SETTING, STORE_AUDIT_LOGS_S3_SETTING, TEAMS_SETTING,
|
||||
TIMEOUT_WAIT_RESULT_SETTING, UV_EXCLUDE_NEWER_SETTING, UV_INDEX_STRATEGY_SETTING,
|
||||
UV_PYTHON_INSTALL_MIRROR_SETTING, WORKSPACE_FAIRNESS_DURATION_SECS_SETTING,
|
||||
WORKSPACE_FAIRNESS_ENABLED_SETTING, WORKSPACE_FAIRNESS_MAX_PERCENT_SETTING,
|
||||
WORKSPACE_FAIRNESS_MIN_TOTAL_SETTING, WORKSPACE_REGISTRIES_SETTING,
|
||||
SANDBOX_IMAGE_CACHE_MAX_MB_SETTING, SANDBOX_IMAGE_DEFAULT_REGISTRY_SETTING,
|
||||
SANDBOX_IMAGE_MAX_SIZE_MB_SETTING, SANDBOX_IMAGE_PULL_POLICY_SETTING,
|
||||
SANDBOX_REGISTRY_AUTH_SETTING, SCIM_TOKEN_SETTING, SMTP_SETTING,
|
||||
STORE_AUDIT_LOGS_S3_SETTING, TEAMS_SETTING, TIMEOUT_WAIT_RESULT_SETTING,
|
||||
UV_EXCLUDE_NEWER_SETTING, UV_INDEX_STRATEGY_SETTING, UV_PYTHON_INSTALL_MIRROR_SETTING,
|
||||
WORKSPACE_FAIRNESS_DURATION_SECS_SETTING, WORKSPACE_FAIRNESS_ENABLED_SETTING,
|
||||
WORKSPACE_FAIRNESS_MAX_PERCENT_SETTING, WORKSPACE_FAIRNESS_MIN_TOTAL_SETTING,
|
||||
WORKSPACE_REGISTRIES_SETTING,
|
||||
},
|
||||
scripts::ScriptLang,
|
||||
stats_oss::schedule_stats,
|
||||
@@ -134,8 +137,11 @@ use crate::monitor::{
|
||||
reload_job_default_timeout_setting, reload_job_isolation_setting, reload_jwt_secret_setting,
|
||||
reload_license_key, reload_npm_config_registry_setting, reload_nsjail_tmp_backing_setting,
|
||||
reload_nsjail_tmpfs_size_setting, reload_otel_tracing_proxy_setting,
|
||||
reload_pip_index_url_setting, reload_retention_period_setting, reload_scim_token_setting,
|
||||
reload_smtp_config, reload_store_audit_logs_s3_setting, reload_uv_exclude_newer_setting,
|
||||
reload_pip_index_url_setting, reload_retention_period_setting,
|
||||
reload_sandbox_image_cache_max_setting, reload_sandbox_image_default_registry_setting,
|
||||
reload_sandbox_image_max_size_setting, reload_sandbox_image_pull_policy_setting,
|
||||
reload_sandbox_registry_auth_setting, reload_scim_token_setting, reload_smtp_config,
|
||||
reload_store_audit_logs_s3_setting, reload_uv_exclude_newer_setting,
|
||||
reload_uv_index_strategy_setting, reload_uv_python_install_mirror_setting,
|
||||
reload_worker_config, MonitorIteration,
|
||||
};
|
||||
@@ -1827,6 +1833,19 @@ async fn process_notify_event(
|
||||
JOB_ISOLATION_SETTING => reload_job_isolation_setting(conn).await,
|
||||
NSJAIL_TMPFS_SIZE_MB_SETTING => reload_nsjail_tmpfs_size_setting(conn).await,
|
||||
NSJAIL_TMP_BACKING_SETTING => reload_nsjail_tmp_backing_setting(conn).await,
|
||||
SANDBOX_IMAGE_MAX_SIZE_MB_SETTING => {
|
||||
reload_sandbox_image_max_size_setting(conn).await
|
||||
}
|
||||
SANDBOX_IMAGE_CACHE_MAX_MB_SETTING => {
|
||||
reload_sandbox_image_cache_max_setting(conn).await
|
||||
}
|
||||
SANDBOX_IMAGE_PULL_POLICY_SETTING => {
|
||||
reload_sandbox_image_pull_policy_setting(conn).await
|
||||
}
|
||||
SANDBOX_IMAGE_DEFAULT_REGISTRY_SETTING => {
|
||||
reload_sandbox_image_default_registry_setting(conn).await
|
||||
}
|
||||
SANDBOX_REGISTRY_AUTH_SETTING => reload_sandbox_registry_auth_setting(conn).await,
|
||||
#[cfg(feature = "parquet")]
|
||||
OBJECT_STORE_CONFIG_SETTING => {
|
||||
if !disable_s3_store {
|
||||
|
||||
+72
-3
@@ -66,7 +66,9 @@ use windmill_common::{
|
||||
OTEL_SETTING, OTEL_TRACING_PROXY_SETTING, PIP_INDEX_URL_SETTING,
|
||||
POWERSHELL_REPO_PAT_SETTING, POWERSHELL_REPO_URL_SETTING, PREVIEW_TAGS_OVERRIDE_SETTING,
|
||||
REQUEST_SIZE_LIMIT_SETTING, REQUIRE_PREEXISTING_USER_FOR_OAUTH_SETTING,
|
||||
RETENTION_PERIOD_SECS_SETTING, SAML_METADATA_SETTING, SCIM_TOKEN_SETTING,
|
||||
RETENTION_PERIOD_SECS_SETTING, SAML_METADATA_SETTING, SANDBOX_IMAGE_CACHE_MAX_MB_SETTING,
|
||||
SANDBOX_IMAGE_DEFAULT_REGISTRY_SETTING, SANDBOX_IMAGE_MAX_SIZE_MB_SETTING,
|
||||
SANDBOX_IMAGE_PULL_POLICY_SETTING, SANDBOX_REGISTRY_AUTH_SETTING, SCIM_TOKEN_SETTING,
|
||||
STORE_AUDIT_LOGS_S3_SETTING, TIMEOUT_WAIT_RESULT_SETTING, UV_EXCLUDE_NEWER_SETTING,
|
||||
UV_INDEX_STRATEGY_SETTING, UV_PYTHON_INSTALL_MIRROR_SETTING,
|
||||
WORKSPACE_FAIRNESS_DURATION_SECS_SETTING, WORKSPACE_FAIRNESS_ENABLED_SETTING,
|
||||
@@ -112,8 +114,10 @@ use windmill_worker::{
|
||||
JOB_DEFAULT_TIMEOUT, JOB_ISOLATION, KEEP_JOB_DIR, MAVEN_REPOS, MAVEN_SETTINGS_XML,
|
||||
NO_DEFAULT_MAVEN, NPMRC, NPM_CONFIG_REGISTRY, NSJAIL_AVAILABLE, NSJAIL_TMPFS_SIZE_MB,
|
||||
NSJAIL_TMP_BACKING, NUGET_CONFIG, OTEL_TRACING_PROXY_SETTINGS, PIP_EXTRA_INDEX_URL,
|
||||
PIP_INDEX_URL, POWERSHELL_REPO_PAT, POWERSHELL_REPO_URL, UNSHARE_PATH, UV_EXCLUDE_NEWER,
|
||||
UV_INDEX_STRATEGY, UV_PYTHON_INSTALL_MIRROR, WORKSPACE_REGISTRIES,
|
||||
PIP_INDEX_URL, POWERSHELL_REPO_PAT, POWERSHELL_REPO_URL, SANDBOX_IMAGE_CACHE_MAX_MB,
|
||||
SANDBOX_IMAGE_DEFAULT_REGISTRY, SANDBOX_IMAGE_MAX_SIZE_MB, SANDBOX_IMAGE_PULL_POLICY,
|
||||
SANDBOX_REGISTRY_AUTH, UNSHARE_PATH, UV_EXCLUDE_NEWER, UV_INDEX_STRATEGY,
|
||||
UV_PYTHON_INSTALL_MIRROR, WORKSPACE_REGISTRIES,
|
||||
};
|
||||
|
||||
#[cfg(feature = "parquet")]
|
||||
@@ -407,6 +411,11 @@ pub async fn initial_load(
|
||||
reload_job_isolation_setting(&conn).await;
|
||||
reload_nsjail_tmpfs_size_setting(&conn).await;
|
||||
reload_nsjail_tmp_backing_setting(&conn).await;
|
||||
reload_sandbox_image_max_size_setting(&conn).await;
|
||||
reload_sandbox_image_cache_max_setting(&conn).await;
|
||||
reload_sandbox_image_pull_policy_setting(&conn).await;
|
||||
reload_sandbox_image_default_registry_setting(&conn).await;
|
||||
reload_sandbox_registry_auth_setting(&conn).await;
|
||||
reload_extra_pip_index_url_setting(&conn).await;
|
||||
reload_pip_index_url_setting(&conn).await;
|
||||
reload_uv_index_strategy_setting(&conn).await;
|
||||
@@ -2045,6 +2054,66 @@ pub async fn reload_nsjail_tmp_backing_setting(conn: &Connection) {
|
||||
.await;
|
||||
}
|
||||
|
||||
pub async fn reload_sandbox_image_max_size_setting(conn: &Connection) {
|
||||
reload_option_setting_with_tracing(
|
||||
conn,
|
||||
SANDBOX_IMAGE_MAX_SIZE_MB_SETTING,
|
||||
"SANDBOX_IMAGE_MAX_SIZE_MB",
|
||||
SANDBOX_IMAGE_MAX_SIZE_MB.clone(),
|
||||
)
|
||||
.await;
|
||||
}
|
||||
|
||||
pub async fn reload_sandbox_image_cache_max_setting(conn: &Connection) {
|
||||
reload_option_setting_with_tracing(
|
||||
conn,
|
||||
SANDBOX_IMAGE_CACHE_MAX_MB_SETTING,
|
||||
"SANDBOX_IMAGE_CACHE_MAX_MB",
|
||||
SANDBOX_IMAGE_CACHE_MAX_MB.clone(),
|
||||
)
|
||||
.await;
|
||||
}
|
||||
|
||||
pub async fn reload_sandbox_image_pull_policy_setting(conn: &Connection) {
|
||||
reload_option_setting_with_tracing(
|
||||
conn,
|
||||
SANDBOX_IMAGE_PULL_POLICY_SETTING,
|
||||
"SANDBOX_IMAGE_PULL_POLICY",
|
||||
SANDBOX_IMAGE_PULL_POLICY.clone(),
|
||||
)
|
||||
.await;
|
||||
}
|
||||
|
||||
pub async fn reload_sandbox_image_default_registry_setting(conn: &Connection) {
|
||||
reload_option_setting_with_tracing(
|
||||
conn,
|
||||
SANDBOX_IMAGE_DEFAULT_REGISTRY_SETTING,
|
||||
"SANDBOX_IMAGE_DEFAULT_REGISTRY",
|
||||
SANDBOX_IMAGE_DEFAULT_REGISTRY.clone(),
|
||||
)
|
||||
.await;
|
||||
}
|
||||
|
||||
pub async fn reload_sandbox_registry_auth_setting(conn: &Connection) {
|
||||
// Secret-aware: the value is a raw docker/podman auth.json with credentials, so
|
||||
// it must never be logged. Load directly (the generic reload_option_setting path
|
||||
// logs the value via load_option_setting_value) and only log a redacted message.
|
||||
let q =
|
||||
match load_value_from_global_settings_with_conn(conn, SANDBOX_REGISTRY_AUTH_SETTING, true)
|
||||
.await
|
||||
{
|
||||
Ok(q) => q,
|
||||
Err(e) => {
|
||||
tracing::error!("Error reloading setting SANDBOX_REGISTRY_AUTH: {e:?}");
|
||||
return;
|
||||
}
|
||||
};
|
||||
let value = q.and_then(|q| serde_json::from_value::<String>(q).ok());
|
||||
let configured = value.as_ref().is_some_and(|v| !v.trim().is_empty());
|
||||
*SANDBOX_REGISTRY_AUTH.write().await = value;
|
||||
tracing::info!("Loaded setting SANDBOX_REGISTRY_AUTH (redacted), configured={configured}");
|
||||
}
|
||||
|
||||
pub async fn reload_job_isolation_setting(conn: &Connection) {
|
||||
let value =
|
||||
match load_value_from_global_settings_with_conn(conn, JOB_ISOLATION_SETTING, true).await {
|
||||
|
||||
@@ -58,6 +58,11 @@ pub const NSJAIL_TMPFS_SIZE_MB_SETTING: &str = "nsjail_tmpfs_size_mb";
|
||||
pub const NSJAIL_TMP_BACKING_SETTING: &str = "nsjail_tmp_backing";
|
||||
pub const NSJAIL_TMP_BACKING_DISK: &str = "disk";
|
||||
pub const NSJAIL_TMP_BACKING_TMPFS: &str = "tmpfs";
|
||||
pub const SANDBOX_IMAGE_MAX_SIZE_MB_SETTING: &str = "sandbox_image_max_size_mb";
|
||||
pub const SANDBOX_IMAGE_CACHE_MAX_MB_SETTING: &str = "sandbox_image_cache_max_mb";
|
||||
pub const SANDBOX_IMAGE_PULL_POLICY_SETTING: &str = "sandbox_image_pull_policy";
|
||||
pub const SANDBOX_IMAGE_DEFAULT_REGISTRY_SETTING: &str = "sandbox_image_default_registry";
|
||||
pub const SANDBOX_REGISTRY_AUTH_SETTING: &str = "sandbox_registry_auth";
|
||||
pub const OBJECT_STORE_CONFIG_SETTING: &str = "object_store_cache_config";
|
||||
pub const HUB_API_SECRET_SETTING: &str = "hub_api_secret";
|
||||
|
||||
|
||||
@@ -976,6 +976,7 @@ const SENSITIVE_SETTINGS: &[&str] = &[
|
||||
"ruby_repos",
|
||||
"powershell_repo_pat",
|
||||
"workspace_registries",
|
||||
"sandbox_registry_auth",
|
||||
];
|
||||
|
||||
/// Object-valued settings that contain sensitive sub-fields.
|
||||
|
||||
@@ -859,6 +859,37 @@ pub struct BashAnnotations {
|
||||
pub sandbox: bool,
|
||||
}
|
||||
|
||||
impl BashAnnotations {
|
||||
/// If the script declares `# sandbox <image>` (an image ref after the sandbox
|
||||
/// annotation), returns that image ref. This selects the daemonless, sandboxed
|
||||
/// container runtime: extract the image's rootfs and run it inside the job's
|
||||
/// nsjail sandbox.
|
||||
///
|
||||
/// A bare `# sandbox` (no image argument) returns `None` and keeps the plain
|
||||
/// nsjail-sandboxed-bash behavior (the `sandbox` boolean modifier). `# docker`
|
||||
/// is unaffected and keeps the legacy v1 (dind/daemon) path.
|
||||
pub fn sandbox_image(code: &str) -> Option<String> {
|
||||
for line in code.lines() {
|
||||
let line = line.trim();
|
||||
if line.is_empty() {
|
||||
continue;
|
||||
}
|
||||
// Mirror the annotation parser: stop at the first non-comment line.
|
||||
if !line.starts_with('#') {
|
||||
break;
|
||||
}
|
||||
let mut tokens = line[1..].split_whitespace();
|
||||
if tokens.next() == Some("sandbox") {
|
||||
// `# sandbox <image>` -> container; bare `# sandbox` -> nsjail bash.
|
||||
if let Some(image) = tokens.next() {
|
||||
return Some(image.to_string());
|
||||
}
|
||||
}
|
||||
}
|
||||
None
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub enum SqlResultCollectionStrategy {
|
||||
LastStatementAllRows,
|
||||
@@ -2224,6 +2255,34 @@ mod tests {
|
||||
use super::*;
|
||||
use std::collections::HashMap;
|
||||
|
||||
#[test]
|
||||
fn test_bash_sandbox_image_annotation() {
|
||||
// `# sandbox <image>` selects the container runtime and returns the image.
|
||||
assert_eq!(
|
||||
BashAnnotations::sandbox_image("# sandbox alpine:latest\necho hi"),
|
||||
Some("alpine:latest".to_string())
|
||||
);
|
||||
// Extra whitespace and a leading non-spaced `#` still work.
|
||||
assert_eq!(
|
||||
BashAnnotations::sandbox_image("#sandbox python:3.12-slim\n"),
|
||||
Some("python:3.12-slim".to_string())
|
||||
);
|
||||
// A bare `# sandbox` (no image) keeps the nsjail-bash modifier -> None.
|
||||
assert_eq!(BashAnnotations::sandbox_image("# sandbox\necho hi"), None);
|
||||
// `sandbox` must be its own token, not a prefix.
|
||||
assert_eq!(BashAnnotations::sandbox_image("# sandboxed foo"), None);
|
||||
// Stops at the first non-comment line (image declared too late is ignored).
|
||||
assert_eq!(
|
||||
BashAnnotations::sandbox_image("echo hi\n# sandbox alpine"),
|
||||
None
|
||||
);
|
||||
// `# docker` is a different annotation -> not a sandbox image.
|
||||
assert_eq!(
|
||||
BashAnnotations::sandbox_image("# docker alpine\necho hi"),
|
||||
None
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_mixed_tags() {
|
||||
let input = vec![
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
name: "docker v2 run"
|
||||
|
||||
mode: ONCE
|
||||
hostname: "container"
|
||||
log_level: ERROR
|
||||
time_limit: {TIMEOUT}
|
||||
|
||||
disable_rl: true
|
||||
|
||||
cwd: {WORKDIR}
|
||||
|
||||
clone_newnet: false
|
||||
clone_newuser: {CLONE_NEWUSER}
|
||||
|
||||
skip_setsid: true
|
||||
keep_caps: false
|
||||
# keep_env forwards nsjail's OWN process env (only windmill-trusted keys: reserved
|
||||
# vars + proxy) to the child. The image's attacker-controlled Env is delivered via
|
||||
# the envar directives below — NEVER nsjail's process env, so a hostile image cannot
|
||||
# set LD_PRELOAD/LD_LIBRARY_PATH/LD_AUDIT on the nsjail binary itself.
|
||||
keep_env: true
|
||||
mount_proc: true
|
||||
|
||||
# Image Env (+ PATH/HOME fallbacks), proto-escaped. Applied to the child only.
|
||||
{ENVARS}
|
||||
|
||||
# Map uid/gid 0 inside the jail to the (single) worker user outside. The image's
|
||||
# rootfs is extracted as the worker user, so a root process inside the container
|
||||
# owns the rootfs and runs like a normal "root in container" — without any subuid
|
||||
# range. Multi-uid images are a later enhancement (newuidmap range).
|
||||
uidmap {
|
||||
inside_id: "0"
|
||||
outside_id: ""
|
||||
count: 1
|
||||
}
|
||||
gidmap {
|
||||
inside_id: "0"
|
||||
outside_id: ""
|
||||
count: 1
|
||||
}
|
||||
|
||||
# The image's root filesystem, bound one top-level entry at a time. Binding the
|
||||
# whole rootfs at "/" trips nsjail's read-only remount of its base root in a
|
||||
# rootless userns ("mount(... MS_REMOUNT|MS_BIND|MS_RDONLY): Operation not
|
||||
# permitted"); per-entry binds sit as rw submounts under nsjail's own tmpfs root
|
||||
# and avoid it. Generated from the extracted rootfs.
|
||||
{ROOTFS_MOUNTS}
|
||||
|
||||
# Pseudo-filesystems the image expects. /tmp honors the same instance settings as
|
||||
# every other nsjail job (nsjail_tmp_backing tmpfs/disk, nsjail_tmpfs_size_mb);
|
||||
# /dev gets the standard nodes; /proc comes from mount_proc (the jail's own pid ns).
|
||||
{TMP_MOUNT_BLOCK}
|
||||
|
||||
mount {
|
||||
src: "/dev/null"
|
||||
dst: "/dev/null"
|
||||
is_bind: true
|
||||
rw: true
|
||||
}
|
||||
|
||||
mount {
|
||||
src: "/dev/zero"
|
||||
dst: "/dev/zero"
|
||||
is_bind: true
|
||||
rw: true
|
||||
}
|
||||
|
||||
mount {
|
||||
src: "/dev/random"
|
||||
dst: "/dev/random"
|
||||
is_bind: true
|
||||
}
|
||||
|
||||
mount {
|
||||
src: "/dev/urandom"
|
||||
dst: "/dev/urandom"
|
||||
is_bind: true
|
||||
}
|
||||
|
||||
# Host DNS config layered over the image's /etc so name resolution works on the
|
||||
# job's network (mandatory:false: some minimal images have no /etc files to shadow).
|
||||
mount {
|
||||
src: "/etc/resolv.conf"
|
||||
dst: "/etc/resolv.conf"
|
||||
is_bind: true
|
||||
mandatory: false
|
||||
}
|
||||
|
||||
mount {
|
||||
src: "/etc/hosts"
|
||||
dst: "/etc/hosts"
|
||||
is_bind: true
|
||||
mandatory: false
|
||||
}
|
||||
|
||||
# `# volume` mounts (and the same-worker /tmp/shared folder). Placed after the
|
||||
# rootfs binds and the tmpfs /tmp so a volume target overrides any colliding image
|
||||
# path and isn't shadowed by the tmpfs. Empty when there are no volumes.
|
||||
{SHARED_MOUNT}
|
||||
|
||||
iface_no_lo: true
|
||||
|
||||
#{DEV}
|
||||
@@ -40,9 +40,9 @@ use crate::handle_child::run_future_with_polling_update_job_poller;
|
||||
|
||||
use crate::{
|
||||
common::{
|
||||
build_args_map, build_command_with_isolation, get_reserved_variables, read_file,
|
||||
read_file_content, resolve_nsjail_timeout, resolve_nsjail_tmp_mount_block, start_child_process,
|
||||
OccupancyMetrics, DEV_CONF_NSJAIL,
|
||||
build_args_map, build_command_with_isolation, get_reserved_variables, raw_to_string,
|
||||
read_file, read_file_content, resolve_nsjail_timeout, resolve_nsjail_tmp_mount_block,
|
||||
start_child_process, OccupancyMetrics, DEV_CONF_NSJAIL,
|
||||
},
|
||||
get_proxy_envs_for_lang,
|
||||
handle_child::handle_child,
|
||||
@@ -57,14 +57,6 @@ lazy_static::lazy_static! {
|
||||
pub static ref ANSI_ESCAPE_RE: Regex = Regex::new(r"\x1b\[[0-9;]*m").unwrap();
|
||||
}
|
||||
|
||||
fn raw_to_string(x: &str) -> String {
|
||||
match serde_json::from_str::<serde_json::Value>(x) {
|
||||
Ok(serde_json::Value::String(x)) => x,
|
||||
Ok(x) => serde_json::to_string(&x).unwrap_or_else(|_| String::new()),
|
||||
_ => String::new(),
|
||||
}
|
||||
}
|
||||
|
||||
#[tracing::instrument(level = "trace", skip_all)]
|
||||
pub async fn handle_bash_job(
|
||||
mem_peak: &mut i32,
|
||||
@@ -84,6 +76,28 @@ pub async fn handle_bash_job(
|
||||
) -> Result<Box<RawValue>, Error> {
|
||||
let annotation = windmill_common::worker::BashAnnotations::parse(&content);
|
||||
|
||||
// `# sandbox <image>` selects the daemonless, nsjail-sandboxed container runtime
|
||||
// (extract the image's rootfs + run it inside the job's sandbox). A bare
|
||||
// `# sandbox` keeps the plain nsjail-bash modifier; `# docker` keeps v1 (dind).
|
||||
if let Some(image) = windmill_common::worker::BashAnnotations::sandbox_image(content) {
|
||||
return crate::docker_v2::handle_docker_v2_job(
|
||||
&image,
|
||||
mem_peak,
|
||||
canceled_by,
|
||||
job,
|
||||
conn,
|
||||
client,
|
||||
parent_runnable_path,
|
||||
content,
|
||||
job_dir,
|
||||
shared_mount,
|
||||
base_internal_url,
|
||||
worker_name,
|
||||
occupancy_metrics,
|
||||
)
|
||||
.await;
|
||||
}
|
||||
|
||||
// Check if sandbox annotation is used but nsjail is not available
|
||||
if annotation.sandbox && NSJAIL_AVAILABLE.is_none() {
|
||||
return Err(Error::ExecutionErr(
|
||||
|
||||
@@ -68,6 +68,16 @@ mount {
|
||||
#[cfg(not(debug_assertions))]
|
||||
pub const DEV_CONF_NSJAIL: &str = "";
|
||||
|
||||
/// Turn a JSON value into the string a shell/CLI arg should receive: a JSON string
|
||||
/// becomes its inner value, anything else is re-serialized compactly.
|
||||
pub(crate) fn raw_to_string(x: &str) -> String {
|
||||
match serde_json::from_str::<serde_json::Value>(x) {
|
||||
Ok(serde_json::Value::String(x)) => x,
|
||||
Ok(x) => serde_json::to_string(&x).unwrap_or_else(|_| String::new()),
|
||||
_ => String::new(),
|
||||
}
|
||||
}
|
||||
|
||||
pub async fn build_args_map<'a>(
|
||||
job: &'a MiniPulledJob,
|
||||
client: &AuthedClient,
|
||||
|
||||
@@ -0,0 +1,683 @@
|
||||
//! Sandboxed container runtime: run a container as a sandboxed subprogram of the job.
|
||||
//!
|
||||
//! Unlike the legacy `# docker` (dind/daemon) path, this has no daemon and no Docker
|
||||
//! API. It splits *pull* from *run*:
|
||||
//!
|
||||
//! 1. **pull/extract** (podman, rootless): materialize the image's root filesystem
|
||||
//! into `{job_dir}/rootfs` and read its OCI config (Env/Cmd/Entrypoint/WorkingDir).
|
||||
//! 2. **run** (the job's own nsjail sandbox): execute the image command with the
|
||||
//! extracted rootfs bound in as the new root, so the container inherits exactly
|
||||
//! the job's confinement (filesystem mask, pid namespace, network, uid) and can't
|
||||
//! escape past what the job itself can reach.
|
||||
//!
|
||||
//! Selected by `# sandbox <image>` (a bare `# sandbox` keeps plain nsjail-bash;
|
||||
//! `# docker` keeps the v1 daemon path). The script body runs inside the image via
|
||||
//! `/bin/sh`; an empty body runs the image's ENTRYPOINT/CMD.
|
||||
|
||||
use std::process::Stdio;
|
||||
|
||||
use serde::Deserialize;
|
||||
use serde_json::{json, value::RawValue};
|
||||
use sqlx::types::Json;
|
||||
use tokio::process::Command;
|
||||
|
||||
use windmill_common::{client::AuthedClient, scripts::ScriptLang};
|
||||
use windmill_common::{
|
||||
error::Error,
|
||||
worker::{to_raw_value, write_file, Connection},
|
||||
};
|
||||
|
||||
use windmill_queue::{append_logs, CanceledBy, MiniPulledJob};
|
||||
|
||||
use crate::{
|
||||
common::{
|
||||
build_args_map, get_reserved_variables, raw_to_string, resolve_nsjail_timeout,
|
||||
resolve_nsjail_tmp_mount_block, start_child_process, OccupancyMetrics, DEV_CONF_NSJAIL,
|
||||
},
|
||||
get_proxy_envs_for_lang,
|
||||
handle_child::handle_child,
|
||||
DISABLE_NUSER, NSJAIL_AVAILABLE, NSJAIL_PATH, SANDBOX_IMAGE_CACHE_MAX_MB,
|
||||
SANDBOX_IMAGE_DEFAULT_REGISTRY, SANDBOX_IMAGE_MAX_SIZE_MB, SANDBOX_IMAGE_PULL_POLICY,
|
||||
SANDBOX_REGISTRY_AUTH,
|
||||
};
|
||||
|
||||
const NSJAIL_CONFIG_RUN_DOCKER_CONTENT: &str = include_str!("../nsjail/run.docker.config.proto");
|
||||
|
||||
const DEFAULT_PATH: &str = "/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin";
|
||||
|
||||
lazy_static::lazy_static! {
|
||||
pub static ref PODMAN_PATH: String =
|
||||
std::env::var("PODMAN_PATH").unwrap_or_else(|_| "podman".to_string());
|
||||
}
|
||||
|
||||
/// Guards against overlapping cache-eviction passes across concurrent jobs.
|
||||
static EVICTION_RUNNING: std::sync::atomic::AtomicBool = std::sync::atomic::AtomicBool::new(false);
|
||||
|
||||
/// podman pull policy from the `sandbox_image_pull_policy` instance setting. `newer`
|
||||
/// (the default when unset/invalid) re-pulls only when the registry digest changed —
|
||||
/// one cheap manifest check per job, no transfer if unchanged — so moving tags like
|
||||
/// `:latest` don't go stale. `missing` is fastest (tags can go stale); `always`
|
||||
/// re-checks every job.
|
||||
async fn pull_policy() -> String {
|
||||
let p = SANDBOX_IMAGE_PULL_POLICY.read().await.clone();
|
||||
match p.as_deref() {
|
||||
Some(p @ ("missing" | "newer" | "always" | "never")) => p.to_string(),
|
||||
_ => "newer".to_string(),
|
||||
}
|
||||
}
|
||||
|
||||
/// `sandbox_image_max_size_mb` instance setting; 0 (or unset/non-positive) = no limit.
|
||||
async fn max_image_size_mb() -> u64 {
|
||||
SANDBOX_IMAGE_MAX_SIZE_MB.read().await.unwrap_or(0).max(0) as u64
|
||||
}
|
||||
|
||||
/// `sandbox_image_cache_max_mb` instance setting; 0 (or unset/non-positive) = unbounded.
|
||||
async fn image_cache_max_mb() -> u64 {
|
||||
SANDBOX_IMAGE_CACHE_MAX_MB.read().await.unwrap_or(0).max(0) as u64
|
||||
}
|
||||
|
||||
/// A ref is registry-qualified if the component before the first `/` looks like a
|
||||
/// host (contains `.` or `:`, or is `localhost`). Bare repos (`alpine`,
|
||||
/// `alpine:latest`, `myorg/img`) are unqualified and resolve against docker.io —
|
||||
/// or the configured default registry.
|
||||
fn registry_qualified(image: &str) -> bool {
|
||||
match image.split_once('/') {
|
||||
None => false,
|
||||
Some((first, _)) => first.contains('.') || first.contains(':') || first == "localhost",
|
||||
}
|
||||
}
|
||||
|
||||
/// Prepend the `sandbox_image_default_registry` instance setting to unqualified image
|
||||
/// refs (fully-qualified refs are left untouched).
|
||||
async fn resolve_image_ref(image: &str) -> String {
|
||||
let registry = SANDBOX_IMAGE_DEFAULT_REGISTRY.read().await.clone();
|
||||
match registry {
|
||||
Some(registry) if !registry.trim().is_empty() && !registry_qualified(image) => {
|
||||
format!("{}/{}", registry.trim().trim_end_matches('/'), image)
|
||||
}
|
||||
_ => image.to_string(),
|
||||
}
|
||||
}
|
||||
|
||||
/// If the `sandbox_registry_auth` instance setting holds a docker/podman `auth.json`
|
||||
/// blob, write it to a per-job authfile (0600, removed with the job) and return its
|
||||
/// path to pass to `podman --authfile`. Returns `None` when unset.
|
||||
async fn write_auth_file(job_dir: &str) -> Result<Option<String>, Error> {
|
||||
let auth = SANDBOX_REGISTRY_AUTH.read().await.clone();
|
||||
let Some(auth) = auth.filter(|a| !a.trim().is_empty()) else {
|
||||
return Ok(None);
|
||||
};
|
||||
let path = format!("{job_dir}/registry_auth.json");
|
||||
// Create 0600 from the start (registry credentials) — no world-readable window.
|
||||
#[cfg(unix)]
|
||||
{
|
||||
use tokio::io::AsyncWriteExt;
|
||||
let mut f = tokio::fs::OpenOptions::new()
|
||||
.write(true)
|
||||
.create(true)
|
||||
.truncate(true)
|
||||
.mode(0o600)
|
||||
.open(&path)
|
||||
.await?;
|
||||
f.write_all(auth.as_bytes()).await?;
|
||||
}
|
||||
#[cfg(not(unix))]
|
||||
tokio::fs::write(&path, auth).await?;
|
||||
Ok(Some(path))
|
||||
}
|
||||
|
||||
/// The subset of an image's OCI config we apply to the run.
|
||||
#[derive(Deserialize, Default, Debug)]
|
||||
struct OciConfig {
|
||||
#[serde(default, rename = "Env")]
|
||||
env: Option<Vec<String>>,
|
||||
#[serde(default, rename = "Cmd")]
|
||||
cmd: Option<Vec<String>>,
|
||||
#[serde(default, rename = "Entrypoint")]
|
||||
entrypoint: Option<Vec<String>>,
|
||||
#[serde(default, rename = "WorkingDir")]
|
||||
working_dir: Option<String>,
|
||||
}
|
||||
|
||||
/// Quote a string as a protobuf-text-format string literal for safe inclusion in
|
||||
/// the nsjail config. Image-controlled values (mount srcs/dsts, symlink targets,
|
||||
/// WorkingDir) flow into the config, so they MUST be escaped — an unescaped `"` or
|
||||
/// newline would otherwise let a hostile image config inject arbitrary nsjail
|
||||
/// directives and break out of the sandbox. Every byte is emitted as a printable
|
||||
/// ASCII char or a valid protobuf escape (`\"`, `\\`, `\n`/`\r`/`\t`, or 3-digit
|
||||
/// octal `\NNN` for control/non-ASCII bytes), so the result always parses.
|
||||
fn proto_str(s: &str) -> String {
|
||||
let mut out = String::with_capacity(s.len() + 2);
|
||||
out.push('"');
|
||||
for &b in s.as_bytes() {
|
||||
match b {
|
||||
b'"' => out.push_str("\\\""),
|
||||
b'\\' => out.push_str("\\\\"),
|
||||
b'\n' => out.push_str("\\n"),
|
||||
b'\r' => out.push_str("\\r"),
|
||||
b'\t' => out.push_str("\\t"),
|
||||
0x20..=0x7e => out.push(b as char),
|
||||
_ => out.push_str(&format!("\\{b:03o}")),
|
||||
}
|
||||
}
|
||||
out.push('"');
|
||||
out
|
||||
}
|
||||
|
||||
/// Render container env vars as nsjail `envar:` directives (one per line). Each
|
||||
/// `KEY=VALUE` is proto-escaped, so image-controlled keys/values can neither break
|
||||
/// the config nor reach nsjail's own process environment.
|
||||
fn render_envars(env: &[(String, String)]) -> String {
|
||||
env.iter()
|
||||
.map(|(k, v)| format!("envar: {}", proto_str(&format!("{k}={v}"))))
|
||||
.collect::<Vec<_>>()
|
||||
.join("\n")
|
||||
}
|
||||
|
||||
async fn podman(args: &[&str]) -> Result<std::process::Output, Error> {
|
||||
Command::new(PODMAN_PATH.as_str())
|
||||
.args(args)
|
||||
.output()
|
||||
.await
|
||||
.map_err(|e| Error::ExecutionErr(format!("failed to run podman {}: {e}", args.join(" "))))
|
||||
}
|
||||
|
||||
/// Pull (if needed) and unpack `image` into `{job_dir}/rootfs`, returning its OCI
|
||||
/// config. Uses podman rootless: `create` (auto-pulls) + `export | tar -x`, with the
|
||||
/// config read from the resulting container (== image config, no command override).
|
||||
async fn extract_image(image: &str, job_dir: &str) -> Result<OciConfig, Error> {
|
||||
let rootfs = format!("{job_dir}/rootfs");
|
||||
tokio::fs::create_dir_all(&rootfs).await?;
|
||||
|
||||
// `podman create` (no command) pulls the image per the configured policy and
|
||||
// records the image's own Cmd/Entrypoint, which we then read back from the
|
||||
// container config. `--` guards against an `image` ref that starts with `-` being
|
||||
// parsed as a flag (e.g. `--authfile=...`) — the ref is attacker-controlled in
|
||||
// the untrusted case.
|
||||
let pull = format!("--pull={}", pull_policy().await);
|
||||
let mut create_args = vec!["create", &pull];
|
||||
let authfile = write_auth_file(job_dir).await?;
|
||||
if let Some(authfile) = authfile.as_deref() {
|
||||
create_args.push("--authfile");
|
||||
create_args.push(authfile);
|
||||
}
|
||||
create_args.push("--");
|
||||
create_args.push(image);
|
||||
let created = podman(&create_args).await?;
|
||||
if !created.status.success() {
|
||||
return Err(Error::ExecutionErr(format!(
|
||||
"failed to pull/create image {image}: {}",
|
||||
String::from_utf8_lossy(&created.stderr)
|
||||
)));
|
||||
}
|
||||
let container_id = String::from_utf8_lossy(&created.stdout).trim().to_string();
|
||||
|
||||
// Always clean up the container, even on a later failure.
|
||||
let result = extract_created(image, &container_id, &rootfs).await;
|
||||
let _ = podman(&["rm", "-f", &container_id]).await;
|
||||
result
|
||||
}
|
||||
|
||||
async fn extract_created(
|
||||
image: &str,
|
||||
container_id: &str,
|
||||
rootfs: &str,
|
||||
) -> Result<OciConfig, Error> {
|
||||
// Reject oversized images before paying the (large) extraction cost.
|
||||
enforce_image_size_limit(image).await?;
|
||||
|
||||
let inspected = podman(&["inspect", container_id, "--format", "{{json .Config}}"]).await?;
|
||||
if !inspected.status.success() {
|
||||
return Err(Error::ExecutionErr(format!(
|
||||
"failed to inspect image {image}: {}",
|
||||
String::from_utf8_lossy(&inspected.stderr)
|
||||
)));
|
||||
}
|
||||
let config: OciConfig = serde_json::from_slice(&inspected.stdout)
|
||||
.map_err(|e| Error::ExecutionErr(format!("failed to parse image {image} config: {e}")))?;
|
||||
|
||||
// Flatten the image's layers into a rootfs directory. Go through a tar on disk
|
||||
// (in the job dir, cleaned up with the job) rather than a shell pipe. Extracted
|
||||
// as the worker user, so the rootfs is owned by the worker user — which the
|
||||
// single-uid jail maps to uid 0 inside.
|
||||
let tar_path = format!("{rootfs}.tar");
|
||||
let exported = podman(&["export", container_id, "--output", &tar_path]).await?;
|
||||
if !exported.status.success() {
|
||||
let _ = tokio::fs::remove_file(&tar_path).await;
|
||||
return Err(Error::ExecutionErr(format!(
|
||||
"failed to export image {image}: {}",
|
||||
String::from_utf8_lossy(&exported.stderr)
|
||||
)));
|
||||
}
|
||||
let untar = Command::new("tar")
|
||||
.args(["-xf", &tar_path, "-C", rootfs])
|
||||
.output()
|
||||
.await
|
||||
.map_err(|e| Error::ExecutionErr(format!("failed to run tar: {e}")))?;
|
||||
let _ = tokio::fs::remove_file(&tar_path).await;
|
||||
if !untar.status.success() {
|
||||
return Err(Error::ExecutionErr(format!(
|
||||
"failed to unpack image {image}: {}",
|
||||
String::from_utf8_lossy(&untar.stderr)
|
||||
)));
|
||||
}
|
||||
|
||||
Ok(config)
|
||||
}
|
||||
|
||||
/// Reject the image if its on-disk (uncompressed) size exceeds
|
||||
/// `SANDBOX_IMAGE_MAX_SIZE_MB`. No-op when the limit is 0 (unset).
|
||||
async fn enforce_image_size_limit(image: &str) -> Result<(), Error> {
|
||||
let max = max_image_size_mb().await;
|
||||
if max == 0 {
|
||||
return Ok(());
|
||||
}
|
||||
let out = podman(&["image", "inspect", image, "--format", "{{.Size}}"]).await?;
|
||||
if !out.status.success() {
|
||||
// Don't silently bypass the guard — surface it so an operator can see the
|
||||
// size limit isn't being enforced for this image.
|
||||
tracing::warn!(
|
||||
"sandbox image size guard: `podman image inspect {image}` failed, not \
|
||||
enforcing SANDBOX_IMAGE_MAX_SIZE_MB: {}",
|
||||
String::from_utf8_lossy(&out.stderr)
|
||||
);
|
||||
return Ok(());
|
||||
}
|
||||
let bytes: u64 = String::from_utf8_lossy(&out.stdout)
|
||||
.trim()
|
||||
.parse()
|
||||
.unwrap_or(0);
|
||||
let mb = bytes / 1_000_000;
|
||||
if mb > max {
|
||||
return Err(Error::ExecutionErr(format!(
|
||||
"image {image} is {mb} MB, over the SANDBOX_IMAGE_MAX_SIZE_MB limit of {max} MB"
|
||||
)));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[derive(Deserialize)]
|
||||
struct PodmanImage {
|
||||
#[serde(rename = "Id")]
|
||||
id: String,
|
||||
// `default`: podman tags Size/Created `omitempty`, so a degenerate image with a
|
||||
// zero value drops the key — without this the whole array would fail to parse.
|
||||
#[serde(default, rename = "Size")]
|
||||
size: u64,
|
||||
#[serde(default, rename = "Created")]
|
||||
created: i64,
|
||||
}
|
||||
|
||||
/// Best-effort eviction: while the summed size of podman's images exceeds
|
||||
/// `SANDBOX_IMAGE_CACHE_MAX_MB`, remove the oldest (by created time, an LRU proxy).
|
||||
/// No-op when the limit is 0 (unset). Skipped if another pass is already running.
|
||||
/// Images currently backing a container (e.g. a concurrent job mid-extract) fail
|
||||
/// `rmi` and stop the pass, so in-use images are never removed.
|
||||
async fn enforce_image_cache_limit() {
|
||||
use std::sync::atomic::Ordering;
|
||||
let max_mb = image_cache_max_mb().await;
|
||||
if max_mb == 0 {
|
||||
return;
|
||||
}
|
||||
if EVICTION_RUNNING
|
||||
.compare_exchange(false, true, Ordering::SeqCst, Ordering::SeqCst)
|
||||
.is_err()
|
||||
{
|
||||
return;
|
||||
}
|
||||
// Reset the guard on every exit path (incl. an early `break` or a panic), so a
|
||||
// stuck flag can never permanently disable eviction until a worker restart.
|
||||
struct ResetOnDrop;
|
||||
impl Drop for ResetOnDrop {
|
||||
fn drop(&mut self) {
|
||||
EVICTION_RUNNING.store(false, std::sync::atomic::Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
let _reset = ResetOnDrop;
|
||||
let max_bytes = max_mb.saturating_mul(1_000_000);
|
||||
loop {
|
||||
let Ok(out) = podman(&["images", "--format", "json"]).await else {
|
||||
break;
|
||||
};
|
||||
if !out.status.success() {
|
||||
break;
|
||||
}
|
||||
let mut imgs: Vec<PodmanImage> = match serde_json::from_slice(&out.stdout) {
|
||||
Ok(v) => v,
|
||||
Err(e) => {
|
||||
// Don't silently disable eviction on a schema hiccup — surface it.
|
||||
tracing::warn!(
|
||||
"sandbox image cache eviction: cannot parse `podman images` json: {e}"
|
||||
);
|
||||
break;
|
||||
}
|
||||
};
|
||||
let total: u64 = imgs.iter().map(|i| i.size).sum();
|
||||
if total <= max_bytes || imgs.is_empty() {
|
||||
break;
|
||||
}
|
||||
imgs.sort_by_key(|i| i.created);
|
||||
let victim = imgs[0].id.clone();
|
||||
match podman(&["rmi", &victim]).await {
|
||||
Ok(rm) if rm.status.success() => {
|
||||
tracing::info!("sandbox image cache eviction: removed {victim}");
|
||||
}
|
||||
Ok(rm) => {
|
||||
tracing::warn!(
|
||||
"sandbox image cache eviction: cannot remove {victim} (in use?): {}",
|
||||
String::from_utf8_lossy(&rm.stderr)
|
||||
);
|
||||
break;
|
||||
}
|
||||
Err(_) => break,
|
||||
}
|
||||
}
|
||||
// `_reset` drops here and clears EVICTION_RUNNING.
|
||||
}
|
||||
|
||||
/// Build the nsjail mount block that binds each top-level entry of the rootfs in
|
||||
/// place. Binding the whole rootfs at `/` trips nsjail's read-only remount of its
|
||||
/// base root in a rootless userns; per-entry binds avoid it. `proc`, `dev`, `tmp`
|
||||
/// and `sys` are skipped — the profile provides them.
|
||||
async fn generate_rootfs_mounts(rootfs: &str) -> Result<String, Error> {
|
||||
let mut block = String::new();
|
||||
let mut entries = tokio::fs::read_dir(rootfs).await?;
|
||||
while let Some(entry) = entries.next_entry().await? {
|
||||
let name = entry.file_name();
|
||||
let name = name.to_string_lossy();
|
||||
if matches!(name.as_ref(), "proc" | "dev" | "tmp" | "sys") {
|
||||
continue;
|
||||
}
|
||||
let src = proto_str(&format!("{rootfs}/{name}"));
|
||||
let dst = proto_str(&format!("/{name}"));
|
||||
let file_type = entry.file_type().await?;
|
||||
if file_type.is_symlink() {
|
||||
// Recreate top-level symlinks (e.g. usr-merged /bin -> usr/bin) as
|
||||
// symlinks in the jail. The target is image-controlled but only ever
|
||||
// *resolved inside the jail* (against the bound rootfs dirs / jail
|
||||
// pseudo-fs) — there is no host `/` in the jail for it to point at — and
|
||||
// it is escaped via proto_str, so it can neither escape nor inject config.
|
||||
let target = tokio::fs::read_link(entry.path())
|
||||
.await
|
||||
.map(|p| p.to_string_lossy().into_owned())
|
||||
.unwrap_or_default();
|
||||
block.push_str(&format!(
|
||||
"mount {{\n src: {}\n dst: {dst}\n is_symlink: true\n mandatory: false\n}}\n",
|
||||
proto_str(&target),
|
||||
));
|
||||
} else {
|
||||
block.push_str(&format!(
|
||||
"mount {{\n src: {src}\n dst: {dst}\n is_bind: true\n rw: true\n mandatory: false\n}}\n",
|
||||
));
|
||||
}
|
||||
}
|
||||
Ok(block)
|
||||
}
|
||||
|
||||
#[tracing::instrument(level = "trace", skip_all)]
|
||||
pub async fn handle_docker_v2_job(
|
||||
image: &str,
|
||||
mem_peak: &mut i32,
|
||||
canceled_by: &mut Option<CanceledBy>,
|
||||
job: &MiniPulledJob,
|
||||
conn: &Connection,
|
||||
client: &AuthedClient,
|
||||
parent_runnable_path: Option<String>,
|
||||
content: &str,
|
||||
job_dir: &str,
|
||||
shared_mount: &str,
|
||||
base_internal_url: &str,
|
||||
worker_name: &str,
|
||||
occupancy_metrics: &mut OccupancyMetrics,
|
||||
) -> Result<Box<RawValue>, Error> {
|
||||
// The sandboxed container runtime *is* nsjail, so it requires nsjail. (`# docker`
|
||||
// keeps the v1 dind path for non-sandboxed workers.)
|
||||
if NSJAIL_AVAILABLE.is_none() {
|
||||
return Err(Error::ExecutionErr(format!(
|
||||
"`# sandbox {image}` runs the image inside nsjail, which is not available on \
|
||||
this worker. Install nsjail, or use a bare `# docker` (dind) instead."
|
||||
)));
|
||||
}
|
||||
|
||||
// Apply the default-registry instance setting to unqualified refs.
|
||||
let resolved_image = resolve_image_ref(image).await;
|
||||
let image = resolved_image.as_str();
|
||||
|
||||
append_logs(
|
||||
&job.id,
|
||||
&job.workspace_id,
|
||||
format!("\n\n--- SANDBOXED CONTAINER (nsjail) ---\nextracting image {image}...\n"),
|
||||
conn,
|
||||
)
|
||||
.await;
|
||||
|
||||
let config = extract_image(image, job_dir).await?;
|
||||
let rootfs = format!("{job_dir}/rootfs");
|
||||
|
||||
// Best-effort: keep podman's image store under its size cap (overlaps the run).
|
||||
tokio::spawn(enforce_image_cache_limit());
|
||||
|
||||
// Resolve the script args from the bash signature, like the bash executor.
|
||||
let args = build_args_map(job, client, conn).await?.map(Json);
|
||||
let job_args = if args.is_some() {
|
||||
args.as_ref()
|
||||
} else {
|
||||
job.args.as_ref()
|
||||
};
|
||||
let args_owned = windmill_parser_bash::parse_bash_sig(content)?
|
||||
.args
|
||||
.iter()
|
||||
.map(|arg| {
|
||||
job_args
|
||||
.and_then(|x| x.get(&arg.name).map(|x| raw_to_string(x.get())))
|
||||
.unwrap_or_else(String::new)
|
||||
})
|
||||
.collect::<Vec<String>>();
|
||||
|
||||
// The body is everything that isn't a leading `#` annotation/comment line. With
|
||||
// a body we run it via the image's `/bin/sh`; without one we run the image's
|
||||
// ENTRYPOINT + CMD.
|
||||
let has_body = content
|
||||
.lines()
|
||||
.any(|l| !l.trim().is_empty() && !l.trim_start().starts_with('#'));
|
||||
|
||||
let cmd_args: Vec<String> = if has_body {
|
||||
// Pass the body straight to `sh -c` rather than writing a script file into
|
||||
// the image-controlled rootfs: a malicious image could plant that path as a
|
||||
// symlink to a host file and capture the worker's write before nsjail starts
|
||||
// (sandbox-boundary bypass). `sh -c <body> sh <args...>` binds args as $1.. .
|
||||
let mut v = vec![
|
||||
"/bin/sh".to_string(),
|
||||
"-c".to_string(),
|
||||
format!("set -e\n{content}"),
|
||||
"sh".to_string(),
|
||||
];
|
||||
v.extend(args_owned.iter().cloned());
|
||||
v
|
||||
} else {
|
||||
let mut v = config.entrypoint.clone().unwrap_or_default();
|
||||
v.extend(config.cmd.clone().unwrap_or_default());
|
||||
if v.is_empty() {
|
||||
return Err(Error::ExecutionErr(format!(
|
||||
"image {image} has no ENTRYPOINT/CMD and the script body is empty — \
|
||||
nothing to run"
|
||||
)));
|
||||
}
|
||||
v.extend(args_owned.iter().cloned());
|
||||
v
|
||||
};
|
||||
|
||||
let working_dir = config
|
||||
.working_dir
|
||||
.as_deref()
|
||||
.filter(|w| !w.is_empty())
|
||||
.unwrap_or("/");
|
||||
|
||||
// The image's OCI Env is attacker-controlled (BOTH keys and values), so it must
|
||||
// NOT enter the nsjail launcher's own process env: a hostile image could set
|
||||
// LD_PRELOAD / LD_LIBRARY_PATH / LD_AUDIT and have the dynamic loader run code in
|
||||
// the nsjail binary as the worker — outside the jail — before it sandboxes.
|
||||
// Deliver it to the *child only* via proto-escaped `envar:` directives.
|
||||
let mut container_env: Vec<(String, String)> = Vec::new();
|
||||
for kv in config.env.unwrap_or_default() {
|
||||
if let Some((k, v)) = kv.split_once('=') {
|
||||
container_env.push((k.to_string(), v.to_string()));
|
||||
}
|
||||
}
|
||||
if !container_env.iter().any(|(k, _)| k == "PATH") {
|
||||
container_env.push(("PATH".to_string(), DEFAULT_PATH.to_string()));
|
||||
}
|
||||
if !container_env.iter().any(|(k, _)| k == "HOME") {
|
||||
container_env.push(("HOME".to_string(), "/root".to_string()));
|
||||
}
|
||||
let envars = render_envars(&container_env);
|
||||
|
||||
// Render the nsjail profile: dynamic per-entry rootfs binds + image WorkingDir.
|
||||
let nsjail_timeout = resolve_nsjail_timeout(conn, &job.workspace_id, job.id, job.timeout).await;
|
||||
let rootfs_mounts = generate_rootfs_mounts(&rootfs).await?;
|
||||
write_file(
|
||||
job_dir,
|
||||
"run.docker.config.proto",
|
||||
&NSJAIL_CONFIG_RUN_DOCKER_CONTENT
|
||||
.replace("{TIMEOUT}", &nsjail_timeout)
|
||||
.replace("{CLONE_NEWUSER}", &(!*DISABLE_NUSER).to_string())
|
||||
// proto_str-quoted: WorkingDir is image-controlled, must not break out
|
||||
// of the `cwd:` string and inject nsjail directives.
|
||||
.replace("{WORKDIR}", &proto_str(working_dir))
|
||||
.replace("{ROOTFS_MOUNTS}", &rootfs_mounts)
|
||||
.replace(
|
||||
"{TMP_MOUNT_BLOCK}",
|
||||
&resolve_nsjail_tmp_mount_block(job_dir).await,
|
||||
)
|
||||
// `# volume` mounts + same-worker shared folder (empty if none).
|
||||
.replace("{SHARED_MOUNT}", shared_mount)
|
||||
// Image env as `envar:` directives (child-only), so it never touches
|
||||
// nsjail's process env.
|
||||
.replace("{ENVARS}", &envars)
|
||||
.replace("#{DEV}", DEV_CONF_NSJAIL),
|
||||
)?;
|
||||
|
||||
// nsjail's OWN process env: only windmill-trusted keys (reserved vars so
|
||||
// `wmill`/API calls work, + proxy). `keep_env: true` forwards these to the
|
||||
// child. The image env is NOT here — see container_env above.
|
||||
let mut reserved_variables =
|
||||
get_reserved_variables(job, &client.token, conn, parent_runnable_path).await?;
|
||||
reserved_variables.insert("RUST_LOG".to_string(), "info".to_string());
|
||||
reserved_variables.insert(
|
||||
"BASE_INTERNAL_URL".to_string(),
|
||||
base_internal_url.to_string(),
|
||||
);
|
||||
|
||||
let proxy_envs = get_proxy_envs_for_lang(
|
||||
&ScriptLang::Bash,
|
||||
job.kind,
|
||||
&job.id,
|
||||
&job.workspace_id,
|
||||
conn,
|
||||
)
|
||||
.await?;
|
||||
|
||||
let mut nsjail_run_args = vec!["--config", "run.docker.config.proto", "--"];
|
||||
nsjail_run_args.extend(cmd_args.iter().map(|s| s.as_str()));
|
||||
|
||||
let mut nsjail_cmd = Command::new(NSJAIL_PATH.as_str());
|
||||
nsjail_cmd
|
||||
.current_dir(job_dir)
|
||||
.env_clear()
|
||||
.envs(reserved_variables)
|
||||
.envs(proxy_envs)
|
||||
.args(nsjail_run_args)
|
||||
.stdin(Stdio::null())
|
||||
.stdout(Stdio::piped())
|
||||
.stderr(Stdio::piped());
|
||||
let child = start_child_process(nsjail_cmd, NSJAIL_PATH.as_str(), false).await?;
|
||||
|
||||
handle_child(
|
||||
&job.id,
|
||||
conn,
|
||||
mem_peak,
|
||||
canceled_by,
|
||||
child,
|
||||
true,
|
||||
worker_name,
|
||||
&job.workspace_id,
|
||||
"sandboxed container run",
|
||||
job.timeout,
|
||||
true,
|
||||
&mut Some(occupancy_metrics),
|
||||
None,
|
||||
None,
|
||||
)
|
||||
.await?;
|
||||
|
||||
Ok(to_raw_value(&json!(format!(
|
||||
"sandboxed container ({image}) completed successfully"
|
||||
))))
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::{proto_str, registry_qualified, render_envars};
|
||||
|
||||
#[test]
|
||||
fn render_envars_emits_proto_directives() {
|
||||
// Image-controlled env (incl. loader vars) is rendered as `envar:` directives
|
||||
// — i.e. delivered to the child via the config, NOT nsjail's process env, so
|
||||
// it can never set LD_PRELOAD/etc. on the nsjail binary itself.
|
||||
let env = vec![
|
||||
("PATH".to_string(), "/usr/bin".to_string()),
|
||||
("LD_PRELOAD".to_string(), "rootfs/evil.so".to_string()),
|
||||
];
|
||||
let out = render_envars(&env);
|
||||
assert_eq!(
|
||||
out,
|
||||
"envar: \"PATH=/usr/bin\"\nenvar: \"LD_PRELOAD=rootfs/evil.so\""
|
||||
);
|
||||
// A value trying to inject extra directives is escaped, not interpreted.
|
||||
let evil = vec![("X".to_string(), "v\"\nclone_newuser: false".to_string())];
|
||||
let line = render_envars(&evil);
|
||||
assert!(line.starts_with("envar: \""));
|
||||
assert!(!line.contains("\nclone_newuser"));
|
||||
assert!(line.contains("\\n"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn proto_str_escapes_injection() {
|
||||
// Normal paths are just wrapped in quotes.
|
||||
assert_eq!(proto_str("/app"), "\"/app\"");
|
||||
// A `"` is escaped so it cannot close the surrounding string and inject
|
||||
// subsequent nsjail directives — this is what the WorkingDir / mount-src
|
||||
// sandboxing fixes depend on.
|
||||
let malicious = "/x\"\nmount { src: \"/\" dst: \"/host\" is_bind: true }\n#";
|
||||
let escaped = proto_str(malicious);
|
||||
assert!(escaped.starts_with('"') && escaped.ends_with('"'));
|
||||
// No raw quote or newline survives inside the rendered literal.
|
||||
let inner = &escaped[1..escaped.len() - 1];
|
||||
assert!(!inner.contains('\n'));
|
||||
assert!(!inner.contains("\"") || inner.contains("\\\""));
|
||||
assert!(escaped.contains("\\\"")); // the inner quote is backslash-escaped
|
||||
assert!(escaped.contains("\\n")); // the newline is escaped
|
||||
// Control and non-ASCII bytes render as valid 3-digit octal escapes (never
|
||||
// a raw byte or an invalid `\u{..}` that nsjail's parser would reject).
|
||||
assert_eq!(proto_str("a\u{1b}b"), "\"a\\033b\""); // ESC (0x1b)
|
||||
assert_eq!(proto_str("é"), "\"\\303\\251\""); // UTF-8 bytes 0xc3 0xa9
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn registry_qualified_classifies_refs() {
|
||||
// Unqualified: bare repos (with/without tag) and docker.io org/repo.
|
||||
for img in ["alpine", "alpine:latest", "myorg/img", "myorg/img:1.2"] {
|
||||
assert!(!registry_qualified(img), "{img} should be unqualified");
|
||||
}
|
||||
// Qualified: the first path component is a host (has `.`/`:`) or localhost.
|
||||
for img in [
|
||||
"ghcr.io/org/img",
|
||||
"registry.example.com/img:tag",
|
||||
"localhost:5000/img",
|
||||
"localhost/img",
|
||||
"host:5000/a/b",
|
||||
] {
|
||||
assert!(registry_qualified(img), "{img} should be qualified");
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -31,6 +31,7 @@ mod csharp_executor;
|
||||
mod dedicated_worker_ee;
|
||||
mod dedicated_worker_oss;
|
||||
mod deno_executor;
|
||||
mod docker_v2;
|
||||
#[cfg(feature = "duckdb")]
|
||||
mod duckdb_executor;
|
||||
mod global_cache;
|
||||
|
||||
@@ -694,6 +694,27 @@ lazy_static::lazy_static! {
|
||||
/// RAM-backed tmpfs sized by `nsjail_tmpfs_size_mb`.
|
||||
pub static ref NSJAIL_TMP_BACKING: Arc<RwLock<Option<String>>> = Arc::new(RwLock::new(None));
|
||||
|
||||
/// Reject a `# sandbox <image>` whose on-disk size exceeds this many MB, before
|
||||
/// extraction. `None`/non-positive = no limit. (`sandbox_image_max_size_mb`.)
|
||||
pub static ref SANDBOX_IMAGE_MAX_SIZE_MB: Arc<RwLock<Option<i64>>> = Arc::new(RwLock::new(None));
|
||||
|
||||
/// Best-effort cap (MB) on podman's sandbox-image store; oldest images evicted
|
||||
/// after a run when exceeded. `None`/non-positive = unbounded. (`sandbox_image_cache_max_mb`.)
|
||||
pub static ref SANDBOX_IMAGE_CACHE_MAX_MB: Arc<RwLock<Option<i64>>> = Arc::new(RwLock::new(None));
|
||||
|
||||
/// podman pull policy for sandbox images (`missing`/`newer`/`always`/`never`).
|
||||
/// `None`/unrecognized falls back to `newer`. (`sandbox_image_pull_policy`.)
|
||||
pub static ref SANDBOX_IMAGE_PULL_POLICY: Arc<RwLock<Option<String>>> = Arc::new(RwLock::new(None));
|
||||
|
||||
/// If set, unqualified sandbox image refs (e.g. `alpine`) are pulled from this
|
||||
/// registry instead of docker.io. Fully-qualified refs are unaffected.
|
||||
/// (`sandbox_image_default_registry`.)
|
||||
pub static ref SANDBOX_IMAGE_DEFAULT_REGISTRY: Arc<RwLock<Option<String>>> = Arc::new(RwLock::new(None));
|
||||
|
||||
/// Optional docker/podman `auth.json` blob for private registries, written to a
|
||||
/// per-job authfile and passed to `podman --authfile`. (`sandbox_registry_auth`.)
|
||||
pub static ref SANDBOX_REGISTRY_AUTH: Arc<RwLock<Option<String>>> = Arc::new(RwLock::new(None));
|
||||
|
||||
/// Optional mirror URL for `uv python install`. Wires to the `UV_PYTHON_INSTALL_MIRROR`
|
||||
/// env var when forwarded to uv. Can be set via the `UV_PYTHON_INSTALL_MIRROR` env var
|
||||
/// or the `uv_python_install_mirror` instance setting.
|
||||
|
||||
+6
-29
@@ -49,25 +49,6 @@ services:
|
||||
|
||||
logging: *default-logging
|
||||
|
||||
# Docker-in-Docker sidecar: provides an isolated Docker daemon so user scripts
|
||||
# can run containers without accessing the host Docker socket.
|
||||
dind:
|
||||
image: docker:dind
|
||||
privileged: true
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
DOCKER_TLS_CERTDIR: ""
|
||||
volumes:
|
||||
- dind-data:/var/lib/docker
|
||||
expose:
|
||||
- 2375
|
||||
healthcheck:
|
||||
test: ["CMD", "docker", "info"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
logging: *default-logging
|
||||
|
||||
windmill_worker:
|
||||
image: ${WM_IMAGE}
|
||||
pull_policy: always
|
||||
@@ -89,22 +70,19 @@ services:
|
||||
# If running with non-root/non-windmill UID (e.g., user: "1001:1001"),
|
||||
# add: - HOME=/tmp
|
||||
- FAVOR_UNSHARE_PID=true
|
||||
# Connect to the dind sidecar instead of the host Docker socket
|
||||
- DOCKER_HOST=tcp://dind:2375
|
||||
depends_on:
|
||||
db:
|
||||
condition: service_healthy
|
||||
dind:
|
||||
condition: service_healthy
|
||||
# to mount the worker folder to debug, KEEP_JOB_DIR=true and mount /tmp/windmill
|
||||
volumes:
|
||||
- worker_dependency_cache:/tmp/windmill/cache
|
||||
- worker_logs:/tmp/windmill/logs
|
||||
## WARNING: mounting the host Docker socket grants user scripts full access to
|
||||
## the host Docker daemon, enabling host filesystem access and privilege escalation.
|
||||
## Only use this if you fully trust all users who can run scripts.
|
||||
## To use it, remove the DOCKER_HOST env var and dind depends_on above,
|
||||
## and uncomment the line below:
|
||||
## Sandboxed containers (`# sandbox <image>`) run daemonless via podman + nsjail
|
||||
## inside the worker itself — no Docker socket or dind sidecar required.
|
||||
## For the legacy full-compat docker (a bare `# docker`, trusted users only),
|
||||
## mount the host Docker socket by uncommenting the line below. WARNING: this
|
||||
## grants user scripts full access to the host Docker daemon (host filesystem
|
||||
## access and privilege escalation) — only use it if you fully trust all users.
|
||||
# - /var/run/docker.sock:/var/run/docker.sock
|
||||
|
||||
logging: *default-logging
|
||||
@@ -237,4 +215,3 @@ volumes:
|
||||
windmill_index: null
|
||||
lsp_cache: null
|
||||
caddy_data: null
|
||||
dind-data: null
|
||||
|
||||
@@ -0,0 +1,106 @@
|
||||
# Sandboxed container runtime (daemonless docker)
|
||||
|
||||
Windmill bash scripts can run a container image. There are **two** runtimes:
|
||||
|
||||
| | legacy `# docker` | sandboxed `# sandbox <image>` |
|
||||
|---|---|---|
|
||||
| selected by | bare `# docker` | `# sandbox <image>` |
|
||||
| runtime | dind / Docker daemon (bollard, `dind` feature) | daemonless: extract rootfs + nsjail-run |
|
||||
| boundary | separate (daemon outside the jail) | the job's own nsjail sandbox |
|
||||
| nsjail | not provided (trusted-tenant) | **required** — this *is* the sandbox |
|
||||
| safety | trusted-tenant | sandboxed (untrusted-capable) |
|
||||
| compat | full `docker run`/`-d`/API | run-a-command subset |
|
||||
|
||||
The three bash annotations are distinct and don't overload each other:
|
||||
|
||||
- `# docker` → legacy daemon docker (unchanged).
|
||||
- `# sandbox` → run the bash script under nsjail.
|
||||
- `# sandbox <image>` → run that image's command under nsjail (this runtime).
|
||||
|
||||
## Using it
|
||||
|
||||
Put the image ref on a `# sandbox` annotation line; the rest of the script runs
|
||||
**inside** that image:
|
||||
|
||||
```bash
|
||||
# sandbox python:3.12-slim
|
||||
name="$1" # windmill args bind positionally, like any bash script
|
||||
python3 -c "import sys; print('hello', sys.argv[1])" "$name"
|
||||
```
|
||||
|
||||
- The body runs via the image's `/bin/sh -c` (so the image needs a shell).
|
||||
- An **empty** body runs the image's `ENTRYPOINT` + `CMD`.
|
||||
- Windmill args (declared `x="$1"`, …) are appended to the command.
|
||||
- The image's `Env`, `WorkingDir` are applied; the windmill reserved variables
|
||||
(`WM_TOKEN`, `BASE_INTERNAL_URL`, …) are injected so `wmill`/API calls work.
|
||||
|
||||
## How it works
|
||||
|
||||
1. **Pull/extract** (podman, rootless): `podman create --pull=<policy> <image>` +
|
||||
`podman export | tar -x` materializes the image's flattened root filesystem
|
||||
into `{job_dir}/rootfs`, and `podman inspect` reads its OCI config. podman's
|
||||
image store dedups pulls across jobs.
|
||||
2. **Run** (the job's nsjail sandbox): nsjail binds each top-level entry of the
|
||||
rootfs in place (binding the whole rootfs at `/` trips nsjail's read-only
|
||||
remount of its base root in a rootless userns), mounts the standard
|
||||
pseudo-filesystems (`/proc` from the jail's pid namespace, a tmpfs `/tmp`,
|
||||
`/dev` nodes), maps uid/gid 0 inside → the worker user outside, and runs the
|
||||
command. The container *is* the jail.
|
||||
|
||||
```
|
||||
# sandbox <image> ─▶ podman create+export ─▶ {job_dir}/rootfs ─▶ nsjail (chroot rootfs)
|
||||
podman inspect (OCI config) ──────────────────▶ Env / Cmd / WorkingDir
|
||||
```
|
||||
|
||||
Because the run is just the job's own nsjail with the image's filesystem as root,
|
||||
the container inherits exactly the job's confinement:
|
||||
|
||||
- **Filesystem**: only the rootfs + the job's mounts are visible — no host `/`,
|
||||
no other job dirs, no dep cache. There is nothing to bind-mount escape to.
|
||||
- **/proc**: the jail's own pid namespace — the worker and other jobs aren't
|
||||
visible.
|
||||
- **uid**: a single-uid jail — an escape lands as the unprivileged worker user.
|
||||
- **network**: the job's network (same as any bash job).
|
||||
|
||||
## Image storage, freshness & limits
|
||||
|
||||
- **Where pulls live:** podman's rootless graph root (default
|
||||
`$HOME/.local/share/containers/storage`) — persistent, dedups pulls across jobs.
|
||||
The per-job extracted rootfs lives in `{job_dir}/rootfs` and is removed with the
|
||||
job; the transient `rootfs.tar` is removed right after extraction.
|
||||
- **Freshness (`SANDBOX_IMAGE_PULL_POLICY`, default `newer`):** `newer` re-pulls
|
||||
only when the registry digest changed (one cheap manifest check per job, no data
|
||||
transfer if unchanged) — so moving tags like `:latest` don't go stale. `missing`
|
||||
is fastest but tags can go stale; `always` re-checks every job. Pinning a digest
|
||||
(`img@sha256:…`) is immutable and never stale.
|
||||
- **Per-image size cap (`SANDBOX_IMAGE_MAX_SIZE_MB`, default 0 = off):** images
|
||||
whose on-disk size exceeds the cap are rejected before extraction.
|
||||
- **Cache size cap (`SANDBOX_IMAGE_CACHE_MAX_MB`, default 0 = off):** best-effort
|
||||
LRU eviction — after a run, the oldest images are removed until podman's image
|
||||
store is back under the cap. In-use images are never removed.
|
||||
|
||||
## Requirements
|
||||
|
||||
- `podman` (rootless) and `tar` on the worker for image pull/extract.
|
||||
- `nsjail` on the worker — **required**. If nsjail is absent, a `# sandbox <image>`
|
||||
job errors clearly (use a bare `# docker` + a daemon instead).
|
||||
|
||||
## Limitations (by design — daemonless, run-to-completion)
|
||||
|
||||
- No `docker run -d` + later `exec`/`attach`/`logs -f`, no `docker build`,
|
||||
`compose`, swarm, healthchecks.
|
||||
- No arbitrary `-v` host bind mounts, `--privileged`, `--cap-add`, `--device`,
|
||||
host namespace sharing.
|
||||
- Images that drop to a non-root uid or chown to arbitrary uids inside need a
|
||||
subuid **range** in the jail (single-uid only today — follow-up: `newuidmap`
|
||||
range mapping).
|
||||
- The script result is a completion message; capture output via stdout/logs.
|
||||
|
||||
## Follow-ups
|
||||
|
||||
- Content-addressed rootfs cache keyed by image digest (today each job re-exports;
|
||||
podman's image store still dedups the network pull).
|
||||
- Pre-pull size guard via `skopeo` manifest inspection (reject before download).
|
||||
- Subuid-range nsjail variant for multi-uid images.
|
||||
- Per-container isolated networking (slirp/pasta).
|
||||
- Support under the non-nsjail `unshare` isolation mode.
|
||||
@@ -1007,21 +1007,6 @@
|
||||
|
||||
function onScriptLanguageTrigger(lang: 'docker' | 'bunnative' | ScriptLang) {
|
||||
if (lang == 'docker') {
|
||||
if (isCloudHosted()) {
|
||||
sendUserToast(
|
||||
'You cannot use Docker scripts on the multi-tenant platform. Use a dedicated instance or self-host windmill instead.',
|
||||
true,
|
||||
[
|
||||
{
|
||||
label: 'Learn more',
|
||||
callback: () => {
|
||||
window.open('https://www.windmill.dev/docs/advanced/docker', '_blank')
|
||||
}
|
||||
}
|
||||
]
|
||||
)
|
||||
return
|
||||
}
|
||||
template = 'docker'
|
||||
} else if (lang == 'bunnative') {
|
||||
template = 'bunnative'
|
||||
|
||||
@@ -7,8 +7,6 @@
|
||||
import FlowScriptPicker from '../pickers/FlowScriptPicker.svelte'
|
||||
import PickHubScript from '../pickers/PickHubScript.svelte'
|
||||
import WorkspaceScriptPicker from '../pickers/WorkspaceScriptPicker.svelte'
|
||||
import { isCloudHosted } from '$lib/cloud'
|
||||
import { sendUserToast } from '$lib/toast'
|
||||
import ToggleButtonGroup from '$lib/components/common/toggleButton-v2/ToggleButtonGroup.svelte'
|
||||
import ToggleButton from '$lib/components/common/toggleButton-v2/ToggleButton.svelte'
|
||||
import { Check, Code, Zap } from 'lucide-svelte'
|
||||
@@ -259,23 +257,6 @@
|
||||
{label}
|
||||
lang={lang == 'docker' ? 'bash' : lang}
|
||||
on:click={() => {
|
||||
if (lang == 'docker') {
|
||||
if (isCloudHosted()) {
|
||||
sendUserToast(
|
||||
'You cannot use Docker scripts on the multi-tenant platform. Use a dedicated instance or self-host windmill instead.',
|
||||
true,
|
||||
[
|
||||
{
|
||||
label: 'Learn more',
|
||||
callback: () => {
|
||||
window.open('https://www.windmill.dev/docs/advanced/docker', '_blank')
|
||||
}
|
||||
}
|
||||
]
|
||||
)
|
||||
return
|
||||
}
|
||||
}
|
||||
dispatch('new', {
|
||||
language: lang == 'docker' ? 'bash' : lang,
|
||||
kind,
|
||||
|
||||
@@ -4,7 +4,6 @@
|
||||
</script>
|
||||
|
||||
<script lang="ts">
|
||||
import { isCloudHosted } from '$lib/cloud'
|
||||
import { sendUserToast } from '$lib/toast'
|
||||
import FlowScriptPickerQuick from '../pickers/FlowScriptPickerQuick.svelte'
|
||||
import { defaultScriptLanguages, processLangs } from '$lib/scripts'
|
||||
@@ -401,24 +400,6 @@
|
||||
{label}
|
||||
lang={lang == 'docker' ? 'bash' : lang}
|
||||
on:click={() => {
|
||||
if (lang == 'docker') {
|
||||
if (isCloudHosted()) {
|
||||
sendUserToast(
|
||||
'You cannot use Docker scripts on the multi-tenant platform. Use a dedicated instance or self-host windmill instead.',
|
||||
true,
|
||||
[
|
||||
{
|
||||
label: 'Learn more',
|
||||
callback: () => {
|
||||
window.open('https://www.windmill.dev/docs/advanced/docker', '_blank')
|
||||
}
|
||||
}
|
||||
]
|
||||
)
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
dispatch('new', {
|
||||
kind: selectedKind,
|
||||
inlineScript: {
|
||||
|
||||
@@ -268,6 +268,60 @@ export const settings: Record<string, Setting[]> = {
|
||||
placeholder: '800',
|
||||
storage: 'setting'
|
||||
},
|
||||
{
|
||||
label: 'Sandbox image max size (MB)',
|
||||
key: 'sandbox_image_max_size_mb',
|
||||
description:
|
||||
'Reject a <code># sandbox <image></code> whose on-disk size exceeds this many MB, before extraction. Leave empty for no limit.',
|
||||
fieldType: 'number',
|
||||
placeholder: 'no limit',
|
||||
storage: 'setting'
|
||||
},
|
||||
{
|
||||
label: 'Sandbox image cache cap (MB)',
|
||||
key: 'sandbox_image_cache_max_mb',
|
||||
description:
|
||||
"Best-effort cap on the worker's sandbox image store (podman). When exceeded, the oldest images are evicted after a run. Leave empty for unbounded.",
|
||||
fieldType: 'number',
|
||||
placeholder: 'unbounded',
|
||||
storage: 'setting'
|
||||
},
|
||||
{
|
||||
label: 'Sandbox image pull policy',
|
||||
key: 'sandbox_image_pull_policy',
|
||||
description:
|
||||
'When to re-pull a <code># sandbox</code> image. <strong>newer</strong> (default) re-pulls only when the registry digest changed, so moving tags like <code>:latest</code> stay fresh without re-downloading unchanged layers. <strong>missing</strong> pulls only if absent (fastest, tags can go stale). <strong>always</strong> re-checks every job.',
|
||||
fieldType: 'select',
|
||||
storage: 'setting',
|
||||
placeholder: 'newer',
|
||||
defaultValue: () => 'newer',
|
||||
select_items: [
|
||||
{ label: 'Newer (default)', value: 'newer' },
|
||||
{ label: 'Missing', value: 'missing' },
|
||||
{ label: 'Always', value: 'always' },
|
||||
{ label: 'Never', value: 'never' }
|
||||
]
|
||||
},
|
||||
{
|
||||
label: 'Sandbox image default registry',
|
||||
key: 'sandbox_image_default_registry',
|
||||
description:
|
||||
'If set, unqualified <code># sandbox</code> images (e.g. <code>alpine</code>) are pulled from this registry instead of <code>docker.io</code>. Fully-qualified refs (e.g. <code>ghcr.io/org/img</code>) are unaffected. Example: <code>myregistry.example.com</code>.',
|
||||
fieldType: 'text',
|
||||
placeholder: 'docker.io',
|
||||
storage: 'setting'
|
||||
},
|
||||
{
|
||||
label: 'Sandbox registry auth',
|
||||
key: 'sandbox_registry_auth',
|
||||
description:
|
||||
'Credentials for private registries used by <code># sandbox</code> images, in docker/podman <code>auth.json</code> format. Written to a per-job authfile (removed with the job) and passed to <code>podman --authfile</code>.',
|
||||
fieldType: 'codearea',
|
||||
codeAreaLang: 'json',
|
||||
placeholder:
|
||||
'{\n "auths": {\n "myregistry.example.com": {\n "auth": "BASE64(username:password)"\n }\n }\n}',
|
||||
storage: 'setting'
|
||||
},
|
||||
{
|
||||
label: 'Default timeout',
|
||||
key: 'job_default_timeout',
|
||||
|
||||
@@ -1114,21 +1114,18 @@ export const PHP_PREPROCESSOR_MODULE_CODE = `function preprocessor(object $event
|
||||
`
|
||||
|
||||
const DOCKER_INIT_CODE = `# shellcheck shell=bash
|
||||
# docker
|
||||
# The annotation "docker" above is important, it tells windmill that after
|
||||
# the end of the bash script, it should manage the container at id $WM_JOB_ID:
|
||||
# pipe logs, monitor memory usage, kill container if job is cancelled.
|
||||
# sandbox alpine:latest
|
||||
# The "# sandbox <image>" annotation runs this script INSIDE the image above,
|
||||
# sandboxed via nsjail: the image's rootfs is extracted (rootless podman) and the
|
||||
# body runs chrooted in it, inheriting the job's confinement. The body runs with
|
||||
# the image's /bin/sh and windmill args bind positionally as $1, $2, ...
|
||||
# Daemonless — no docker run/-d/exec/build and no host -v bind mounts.
|
||||
# (A bare "# docker" still uses the legacy daemon runtime instead.)
|
||||
|
||||
msg="\${1:-world}"
|
||||
|
||||
IMAGE="alpine:latest"
|
||||
COMMAND="/bin/echo Hello $msg"
|
||||
|
||||
# ensure that the image is up-to-date
|
||||
docker pull $IMAGE
|
||||
|
||||
# if using the 'docker' mode, name it with $WM_JOB_ID for windmill to monitor it
|
||||
docker run --name $WM_JOB_ID -it -d $IMAGE $COMMAND
|
||||
echo "Hello $msg"
|
||||
cat /etc/os-release | head -1
|
||||
`
|
||||
|
||||
const POWERSHELL_INIT_CODE = `param($Msg, [string[]]$Names, [PSCustomObject]$Obj, $Dflt = "default value", [int]$Nb = 3)
|
||||
|
||||
Reference in New Issue
Block a user