Skip to content

bug(docker): delayed exit snapshot regresses restarted sandbox from Ready to Error #4129

Description

@elezar

User Story

As an operator using the Docker compute driver, I want an on-failure sandbox restart to remain Ready once its replacement runtime is healthy, so that delayed observations from the previous run do not block exec or other gateway-managed operations.

Problem Statement

A controlled tmachine reproduction shows that an authentic ContainerExited snapshot from the previous run can regress a newly restarted Docker sandbox from Ready to Error after the replacement supervisor connects. Both replacement containers remain running, with the supervisor healthy.

The final sandbox details show Restart count: 1, no last exit code, and Ready: False (ContainerExited). openshell sandbox exec refuses the sandbox because its phase is Error.

This is a demonstrated stale-state publication failure mode. It is distinct from the Podman workload-containment race in #3776 / PR #4117 and belongs to the cross-driver lifecycle coordination audit in #4115.

The Docker E2E failure has a matching symptom: on_failure_policy_replaces_runtime_and_preserves_workspace timed out after 60 seconds, with the replacement supervisor connected before a Ready -> Error / ContainerExited transition. The controlled reproduction does not establish that the original CI run took this exact path.

Impact / Why This Matters

Gateway-managed exec becomes unavailable despite running replacement containers. The sandbox remains incorrectly marked Error rather than reflecting its live runtime state. In the related CI run, the lifecycle suite finished with 17 passed and 1 failed.

Repeating the workload without forced scheduling passed five times, so retrying may conceal the defect without making restart reliable. No supported recovery workaround was qualified in this investigation; the disposable failed sandbox was deleted after collecting evidence. Recreating a sandbox is an inadequate operational workaround when workspace continuity and a reliable control-plane view are required.

Acceptance Criteria

  • Delivering the previous run's exit snapshot after replacement readiness does not regress a healthy replacement from Ready to Error.
  • After an on-failure restart, exec succeeds and both initial and replacement workspace markers remain readable.
  • A genuine exit of the replacement workload remains observable and is not discarded as stale.
  • Genuine supervisor loss after admission still contains the workload and reports the failure.
  • Deterministic regression coverage forces delayed old-run snapshot delivery across restart/readiness and checks gateway state alongside actual runtime state.
  • Live Docker qualification repeats the controlled overlap and the genuine-loss check; the fix does not rely solely on timing-based stress loops.

Reproduction Steps

The successful reproduction uses temporary scheduling instrumentation. An unmodified live reproduction has not been established.

  1. At checkout ef9adc2f79c2f6161cf4b1625a4c60e8c8958f46, prepare a Docker tmachine guest:

    nix run .#build-artifacts-binaries
    nix run .#build-artifacts-images
    nix run .#tmachine -- setup ubuntu-docker-rootful
    nix run .#tmachine -- test ubuntu-docker-rootful binaries shell
  2. Run this workload in the guest. The initial run creates a persistent marker and exits 17; the replacement writes another marker and stays alive:

    openshell sandbox create --name restart-probe-d-1 \
      --from docker.io/alpine:3.22 \
      --restart-policy on-failure --detach -- sh -lc '
    marker=/sandbox/.openshell-restart-probe
    if [ -e "$marker" ]; then
      echo replacement > /sandbox/replacement-run
      echo replacement-main-ready
      sleep 300
    else
      touch "$marker"
      echo initial > /sandbox/initial-run
      echo initial-main-ready
      sleep 2
      exit 17
    fi'
  3. For the controlled overlap, insert the temporary probe below at the beginning of ComputeRuntime::apply_sandbox_update, before acquiring sync_lock. Rebuild the gateway with nix run .#build-artifacts-binaries, install it in the disposable guest, and set OPENSHELL_DOCKER_RESTART_PROBE=delay in the gateway service environment. Restart the service and wait for openshell status to report Connected before creating the sandbox.

    Temporary A/B instrumentation
    let probe_mode = std::env::var("OPENSHELL_DOCKER_RESTART_PROBE").unwrap_or_default();
    let probe_exit = incoming.name.starts_with("restart-probe-")
        && incoming.status.as_ref().is_some_and(|status| {
            status.conditions.iter().any(|condition| {
                condition.reason == "ContainerExited"
                    && condition.status.eq_ignore_ascii_case("false")
            })
        });
    if probe_exit && !probe_mode.is_empty() {
        info!(sandbox_id = %incoming.id, mode = %probe_mode,
            "RESTART_PROBE holding observed exit snapshot for three seconds");
        tokio::time::sleep(Duration::from_secs(3)).await;
        if probe_mode == "revalidate" {
            if let Ok(Some(live)) = self.get_driver_sandbox(&incoming.id, &incoming.name).await {
                info!(sandbox_id = %incoming.id,
                    "RESTART_PROBE replacing delayed snapshot with live driver state");
                incoming = live;
            }
        }
    }

    This delays an actual incoming exit snapshot; it does not synthesize exit state, stop containers, or delay supervisor startup. It is disposable investigation code, not a production fix.

  4. Poll openshell sandbox get restart-probe-d-1 and capture Docker events and docker inspect state. The replacement reaches Ready, then the held snapshot sets Error / ContainerExited while both containers remain running. Exec rejects the Error phase. Preserve evidence before deleting the test sandbox.

  5. Repeat with OPENSHELL_DOCKER_RESTART_PROBE=revalidate, using a new sandbox name such as restart-probe-r-1. After replacement readiness, observe for at least five seconds and verify:

    openshell sandbox exec restart-probe-r-1 -- \
      cat /sandbox/initial-run /sandbox/replacement-run
  6. Qualify real supervisor loss separately: create a detached sleep 300 sandbox with restart policy never, wait for Ready and another five seconds for driver admission to settle, then kill only that sandbox's supervisor container. Confirm the workload exits and the gateway reports ControlSupervisorExited.

Remove the temporary instrumentation and rebuild clean artifacts after the experiment. The source probe was removed and the disposable guest shut down after this investigation.

Environment

  • Investigation date: 2026-10-02.
  • Gateway: 0.1.3-dev.51+gef9adc2f7.
  • Checkout: ef9adc2f79c2f6161cf4b1625a4c60e8c8958f46; existing Podman/doc edits were preserved. Docker delayed/revalidated trials used only the temporary gateway instrumentation above.
  • Host: macOS ARM64, tmachine/QEMU.
  • Guest: Ubuntu 24.04.4 LTS, ARM64, Linux 6.8.0-138-generic.
  • Runtime: rootful Docker Engine 29.8.2.
  • Environment: ubuntu-docker-rootful, binaries installer; locally built sandbox/supervisor images.
  • Workload: docker.io/alpine:3.22.

Logs and Qualification Results

Experiment Result
Unmodified gateway 5/5 on-failure restarts reached Ready; exec read both workspace markers
Three-second delayed exit snapshot 1/1 regressed Ready to Error / ContainerExited; both replacement containers remained running
Same delay with fresh-state revalidation 3/3 stayed Ready through a five-second observation window; exec read both markers
Genuine supervisor loss after admission, revalidation mode enabled Workload exited; gateway reported ControlSupervisorExited

Controlled delayed trial, times UTC:

15:20:13.249312  Gateway holds an observed ContainerExited snapshot
15:20:13.408960  Replacement workload starts (Docker inspect)
15:20:13.505101  Replacement supervisor starts (Docker inspect)
15:20:13.870166  Gateway logs runtime restarted
15:20:16.251411  Held snapshot is applied: Ready -> Error

Later Docker inspection:

workload:   Status=running Running=true OOMKilled=false
supervisor: Status=running Running=true Health.Status=healthy

CLI:

sandbox 'restart-probe-d-1' is not ready (phase: Error); wait for it to reach Ready state

An initial supervisor-loss attempt killed the supervisor during admission and resulted in ControlSupervisorStartFailed plus cleanup. It is excluded from the settled-runtime containment qualification; the successful check waited another five seconds after Ready.

Investigation and Related Work

Fresh-state revalidation is an experimental A/B comparison. The probe is not serialized against lifecycle mutations and does not establish runtime-generation ownership, cross-process fencing, or cancellation safety. A production fix must coordinate observation validity with lifecycle intent and preserve real replacement failures.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions