User Story
As an operator using the Docker compute driver, I want an on-failure sandbox restart to remain Ready once its replacement runtime is healthy, so that delayed observations from the previous run do not block exec or other gateway-managed operations.
Problem Statement
A controlled tmachine reproduction shows that an authentic ContainerExited snapshot from the previous run can regress a newly restarted Docker sandbox from Ready to Error after the replacement supervisor connects. Both replacement containers remain running, with the supervisor healthy.
The final sandbox details show Restart count: 1, no last exit code, and Ready: False (ContainerExited). openshell sandbox exec refuses the sandbox because its phase is Error.
This is a demonstrated stale-state publication failure mode. It is distinct from the Podman workload-containment race in #3776 / PR #4117 and belongs to the cross-driver lifecycle coordination audit in #4115.
The Docker E2E failure has a matching symptom: on_failure_policy_replaces_runtime_and_preserves_workspace timed out after 60 seconds, with the replacement supervisor connected before a Ready -> Error / ContainerExited transition. The controlled reproduction does not establish that the original CI run took this exact path.
Impact / Why This Matters
Gateway-managed exec becomes unavailable despite running replacement containers. The sandbox remains incorrectly marked Error rather than reflecting its live runtime state. In the related CI run, the lifecycle suite finished with 17 passed and 1 failed.
Repeating the workload without forced scheduling passed five times, so retrying may conceal the defect without making restart reliable. No supported recovery workaround was qualified in this investigation; the disposable failed sandbox was deleted after collecting evidence. Recreating a sandbox is an inadequate operational workaround when workspace continuity and a reliable control-plane view are required.
Acceptance Criteria
Reproduction Steps
The successful reproduction uses temporary scheduling instrumentation. An unmodified live reproduction has not been established.
-
At checkout ef9adc2f79c2f6161cf4b1625a4c60e8c8958f46, prepare a Docker tmachine guest:
nix run .#build-artifacts-binaries
nix run .#build-artifacts-images
nix run .#tmachine -- setup ubuntu-docker-rootful
nix run .#tmachine -- test ubuntu-docker-rootful binaries shell
-
Run this workload in the guest. The initial run creates a persistent marker and exits 17; the replacement writes another marker and stays alive:
openshell sandbox create --name restart-probe-d-1 \
--from docker.io/alpine:3.22 \
--restart-policy on-failure --detach -- sh -lc '
marker=/sandbox/.openshell-restart-probe
if [ -e "$marker" ]; then
echo replacement > /sandbox/replacement-run
echo replacement-main-ready
sleep 300
else
touch "$marker"
echo initial > /sandbox/initial-run
echo initial-main-ready
sleep 2
exit 17
fi'
-
For the controlled overlap, insert the temporary probe below at the beginning of ComputeRuntime::apply_sandbox_update, before acquiring sync_lock. Rebuild the gateway with nix run .#build-artifacts-binaries, install it in the disposable guest, and set OPENSHELL_DOCKER_RESTART_PROBE=delay in the gateway service environment. Restart the service and wait for openshell status to report Connected before creating the sandbox.
Temporary A/B instrumentation
let probe_mode = std::env::var("OPENSHELL_DOCKER_RESTART_PROBE").unwrap_or_default();
let probe_exit = incoming.name.starts_with("restart-probe-")
&& incoming.status.as_ref().is_some_and(|status| {
status.conditions.iter().any(|condition| {
condition.reason == "ContainerExited"
&& condition.status.eq_ignore_ascii_case("false")
})
});
if probe_exit && !probe_mode.is_empty() {
info!(sandbox_id = %incoming.id, mode = %probe_mode,
"RESTART_PROBE holding observed exit snapshot for three seconds");
tokio::time::sleep(Duration::from_secs(3)).await;
if probe_mode == "revalidate" {
if let Ok(Some(live)) = self.get_driver_sandbox(&incoming.id, &incoming.name).await {
info!(sandbox_id = %incoming.id,
"RESTART_PROBE replacing delayed snapshot with live driver state");
incoming = live;
}
}
}
This delays an actual incoming exit snapshot; it does not synthesize exit state, stop containers, or delay supervisor startup. It is disposable investigation code, not a production fix.
-
Poll openshell sandbox get restart-probe-d-1 and capture Docker events and docker inspect state. The replacement reaches Ready, then the held snapshot sets Error / ContainerExited while both containers remain running. Exec rejects the Error phase. Preserve evidence before deleting the test sandbox.
-
Repeat with OPENSHELL_DOCKER_RESTART_PROBE=revalidate, using a new sandbox name such as restart-probe-r-1. After replacement readiness, observe for at least five seconds and verify:
openshell sandbox exec restart-probe-r-1 -- \
cat /sandbox/initial-run /sandbox/replacement-run
-
Qualify real supervisor loss separately: create a detached sleep 300 sandbox with restart policy never, wait for Ready and another five seconds for driver admission to settle, then kill only that sandbox's supervisor container. Confirm the workload exits and the gateway reports ControlSupervisorExited.
Remove the temporary instrumentation and rebuild clean artifacts after the experiment. The source probe was removed and the disposable guest shut down after this investigation.
Environment
- Investigation date: 2026-10-02.
- Gateway:
0.1.3-dev.51+gef9adc2f7.
- Checkout:
ef9adc2f79c2f6161cf4b1625a4c60e8c8958f46; existing Podman/doc edits were preserved. Docker delayed/revalidated trials used only the temporary gateway instrumentation above.
- Host: macOS ARM64, tmachine/QEMU.
- Guest: Ubuntu 24.04.4 LTS, ARM64, Linux
6.8.0-138-generic.
- Runtime: rootful Docker Engine
29.8.2.
- Environment:
ubuntu-docker-rootful, binaries installer; locally built sandbox/supervisor images.
- Workload:
docker.io/alpine:3.22.
Logs and Qualification Results
| Experiment |
Result |
| Unmodified gateway |
5/5 on-failure restarts reached Ready; exec read both workspace markers |
| Three-second delayed exit snapshot |
1/1 regressed Ready to Error / ContainerExited; both replacement containers remained running |
| Same delay with fresh-state revalidation |
3/3 stayed Ready through a five-second observation window; exec read both markers |
| Genuine supervisor loss after admission, revalidation mode enabled |
Workload exited; gateway reported ControlSupervisorExited |
Controlled delayed trial, times UTC:
15:20:13.249312 Gateway holds an observed ContainerExited snapshot
15:20:13.408960 Replacement workload starts (Docker inspect)
15:20:13.505101 Replacement supervisor starts (Docker inspect)
15:20:13.870166 Gateway logs runtime restarted
15:20:16.251411 Held snapshot is applied: Ready -> Error
Later Docker inspection:
workload: Status=running Running=true OOMKilled=false
supervisor: Status=running Running=true Health.Status=healthy
CLI:
sandbox 'restart-probe-d-1' is not ready (phase: Error); wait for it to reach Ready state
An initial supervisor-loss attempt killed the supervisor during admission and resulted in ControlSupervisorStartFailed plus cleanup. It is excluded from the settled-runtime containment qualification; the successful check waited another five seconds after Ready.
Investigation and Related Work
Fresh-state revalidation is an experimental A/B comparison. The probe is not serialized against lifecycle mutations and does not establish runtime-generation ownership, cross-process fencing, or cancellation safety. A production fix must coordinate observation validity with lifecycle intent and preserve real replacement failures.
User Story
As an operator using the Docker compute driver, I want an on-failure sandbox restart to remain Ready once its replacement runtime is healthy, so that delayed observations from the previous run do not block exec or other gateway-managed operations.
Problem Statement
A controlled tmachine reproduction shows that an authentic
ContainerExitedsnapshot from the previous run can regress a newly restarted Docker sandbox fromReadytoErrorafter the replacement supervisor connects. Both replacement containers remain running, with the supervisor healthy.The final sandbox details show
Restart count: 1, no last exit code, andReady: False (ContainerExited).openshell sandbox execrefuses the sandbox because its phase is Error.This is a demonstrated stale-state publication failure mode. It is distinct from the Podman workload-containment race in #3776 / PR #4117 and belongs to the cross-driver lifecycle coordination audit in #4115.
The Docker E2E failure has a matching symptom:
on_failure_policy_replaces_runtime_and_preserves_workspacetimed out after 60 seconds, with the replacement supervisor connected before aReady -> Error / ContainerExitedtransition. The controlled reproduction does not establish that the original CI run took this exact path.Impact / Why This Matters
Gateway-managed exec becomes unavailable despite running replacement containers. The sandbox remains incorrectly marked Error rather than reflecting its live runtime state. In the related CI run, the lifecycle suite finished with 17 passed and 1 failed.
Repeating the workload without forced scheduling passed five times, so retrying may conceal the defect without making restart reliable. No supported recovery workaround was qualified in this investigation; the disposable failed sandbox was deleted after collecting evidence. Recreating a sandbox is an inadequate operational workaround when workspace continuity and a reliable control-plane view are required.
Acceptance Criteria
Reproduction Steps
The successful reproduction uses temporary scheduling instrumentation. An unmodified live reproduction has not been established.
At checkout
ef9adc2f79c2f6161cf4b1625a4c60e8c8958f46, prepare a Docker tmachine guest:Run this workload in the guest. The initial run creates a persistent marker and exits 17; the replacement writes another marker and stays alive:
For the controlled overlap, insert the temporary probe below at the beginning of
ComputeRuntime::apply_sandbox_update, before acquiringsync_lock. Rebuild the gateway withnix run .#build-artifacts-binaries, install it in the disposable guest, and setOPENSHELL_DOCKER_RESTART_PROBE=delayin the gateway service environment. Restart the service and wait foropenshell statusto report Connected before creating the sandbox.Temporary A/B instrumentation
This delays an actual incoming exit snapshot; it does not synthesize exit state, stop containers, or delay supervisor startup. It is disposable investigation code, not a production fix.
Poll
openshell sandbox get restart-probe-d-1and capture Docker events anddocker inspectstate. The replacement reaches Ready, then the held snapshot sets Error / ContainerExited while both containers remain running. Exec rejects the Error phase. Preserve evidence before deleting the test sandbox.Repeat with
OPENSHELL_DOCKER_RESTART_PROBE=revalidate, using a new sandbox name such asrestart-probe-r-1. After replacement readiness, observe for at least five seconds and verify:openshell sandbox exec restart-probe-r-1 -- \ cat /sandbox/initial-run /sandbox/replacement-runQualify real supervisor loss separately: create a detached
sleep 300sandbox with restart policynever, wait for Ready and another five seconds for driver admission to settle, then kill only that sandbox's supervisor container. Confirm the workload exits and the gateway reportsControlSupervisorExited.Remove the temporary instrumentation and rebuild clean artifacts after the experiment. The source probe was removed and the disposable guest shut down after this investigation.
Environment
0.1.3-dev.51+gef9adc2f7.ef9adc2f79c2f6161cf4b1625a4c60e8c8958f46; existing Podman/doc edits were preserved. Docker delayed/revalidated trials used only the temporary gateway instrumentation above.6.8.0-138-generic.29.8.2.ubuntu-docker-rootful,binariesinstaller; locally built sandbox/supervisor images.docker.io/alpine:3.22.Logs and Qualification Results
Controlled delayed trial, times UTC:
Later Docker inspection:
CLI:
An initial supervisor-loss attempt killed the supervisor during admission and resulted in
ControlSupervisorStartFailedplus cleanup. It is excluded from the settled-runtime containment qualification; the successful check waited another five seconds after Ready.Investigation and Related Work
apply_sandbox_updaterevalidates snapshots against live driver state when durable phase is Starting. If the replacement has already reached Ready, it applies the incoming snapshot directly. The probe exercises that branch with an old exit observation.stale_polled_exitfilters stale polled publication, but cannot retract an observation already queued at the gateway.stop_sandboxpublishes a container snapshot throughpublish_container_snapshotafter stopping the runtime.Phase: Errordespite healthy container and working gateway RPC channel #3440 is historical evidence of a similar Error/healthy-container symptom, not an established shared root cause.Fresh-state revalidation is an experimental A/B comparison. The probe is not serialized against lifecycle mutations and does not establish runtime-generation ownership, cross-process fencing, or cancellation safety. A production fix must coordinate observation validity with lifecycle intent and preserve real replacement failures.