Repository navigation
bug(kubernetes): stop supervisor before suspending workload #3421
Description
Activity
- addedarea:sandboxSandbox runtime and isolation workSandbox runtime and isolation worktest:e2e-kubernetesRequires Kubernetes end-to-end coverageRequires Kubernetes end-to-end coverage
on Sep 17, 2026 🏗️ build-plan
Implementation Plan
Issue type:
fix
Complexity: Medium
Confidence: High — the failure and required ordering are clear; the main risk is keeping teardown idempotent under reconciliation races.Summary
Add a crash-resumable
stopping-supervisorphase. Keep the workload boundary reachable while Kubernetes terminates the supervisor, then atomically advance to the existing rollback phase and suspend the workload. Give supervisor deletion its configured Pod grace period plus Kubernetes API observation headroom.Scope
crates/openshell-driver-kubernetes/src/driver.rs: add the durable stop phase, split supervisor Pod deletion from generation Secret cleanup, resume both teardown phases through reconciliation, and calculate the correct deletion deadline.crates/openshell-driver-kubernetes/src/sandbox_runtime.rs: render an explicit supervisor termination grace period and test the Pod contract.architecture/compute-runtimes.md: document that Kubernetes backend cleanup precedes workload suspension and survives gateway restarts.- Temporary quarantine files from PR test(kubernetes): quarantine lifecycle conformance #3423: remove only if that PR merges before this fix.
Implementation Steps
- Render an explicit supervisor Pod termination grace period and calculate deletion time as grace plus
KUBE_API_TIMEOUT, retaining the Kubernetes 30-second fallback for existing Pods. - Add a resource-version-guarded
stopping-supervisorphase that records stop intent without changing the Agent Sandbox desired running state. - Delete the UID-matched supervisor Pod while its workload boundary remains reachable, without deleting generation Secrets yet.
- Advance atomically to
rolling-backandSuspendedorreplicas: 0, then wait for workload and Sandbox stop completion. - Clean generation Secrets and clear lifecycle annotations only after workload teardown completes.
- Teach periodic reconciliation and repeated stop requests to resume either phase safely after interruption.
- Verify the existing
sandbox-lifecyclescenario across shared v1alpha1, shared v1beta1, managed workspace, and external-driver Kubernetes lanes.
Test Plan
- Unit tests: explicit/default supervisor grace, grace-plus-headroom deadline, API-version-specific stop-begin and stop-advance patches, phase parsing, idempotent/resumed transitions, and Pod rendering.
- Integration tests: use an existing mocked Kubernetes API fixture for delayed deletion if one exists; do not introduce a new mock framework solely for this issue.
- E2E tests: run the existing
sandbox-lifecyclescenario in the four affected Kubernetes modes and confirm the supervisor no longer exits 137.
Risks & Open Questions
- Stop RPCs and periodic reconciliation can race, so deletion and transitions must retain UID and resource-version guards.
- A genuinely hung supervisor may use its full grace period; Kubernetes remains the trusted fallback, while the driver waits with API headroom.
- Secret cleanup must remain separate so the RBAC defect in fix(helm): grant secret cleanup permissions #3363 cannot prevent workload suspension.
- PR test(kubernetes): quarantine lifecycle conformance #3423 already contains the temporary quarantine. Coordinate whether this fix supersedes it or follows it.
Documentation Impact
- Update
architecture/compute-runtimes.md. - No gateway TOML, Helm values, public API, or user-facing compute-driver configuration changes.
- No SELinux or AppArmor impact is expected because process identity,
/proc, execution, visibility, and security contexts are unchanged.
Revision 1 — initial plan
- addedstate:acceptedA maintainer decided OpenShell should pursue this issueA maintainer decided OpenShell should pursue this issue
on Sep 17, 2026 🏗️ build-from-issue-agent
Implemented the durable Kubernetes teardown fix in #3424.
The driver now records a crash-resumable
stopping-supervisorphase, removes the supervisor while the workload boundary is still reachable, then suspends the workload and cleans generation Secrets. Supervisor Pod deletion also waits for the configured termination grace period plus Kubernetes API observation headroom.Local verification:
mise exec -- cargo test -p openshell-driver-kubernetes --lib— 224 passedmise run test— passedmise run pre-commit— passed
Kubernetes E2E will run in branch CI because the local host has neither k3d nor a configured Kubernetes context.
- added a commit that references this issue
on Sep 17, 2026 - changed the title
[-]bug(kubernetes): supervisor deletion races the Pod termination grace period[/-][+]bug(kubernetes): stop supervisor before suspending workload[/+]on Sep 17, 2026 - added a commit that references this issue
on Sep 17, 2026 - added a commit that references this issue
on Sep 18, 2026
User Story
As an OpenShell operator using the Kubernetes compute driver, I want sandbox stop and start operations to complete reliably, so that lifecycle operations work across supported Agent Sandbox API versions and workspace modes.
Problem Statement
Kubernetes
sandbox stopused the wrong teardown order for the separately scheduled supervisor introduced by RFC 0012:timed out waiting for supervisor Pod deletion.This is a deterministic sequencing defect, not a timing race between the driver timeout and the Pod termination grace period. The matching 30-second values made the failure present as a deletion-timeout race, but increasing either timeout alone would not restore the boundary required for normal supervisor shutdown.
The correct order is to record durable stop intent, delete the supervisor while the workload boundary remains reachable, wait for supervisor deletion, and only then suspend the workload. The durable phase is required so reconciliation can resume either step after interruption without reversing the order.
The failure is exercised by the
sandbox-lifecycleconformance scenario added in #3375. Generation Secret cleanup also encountered a separate Helm RBAC mismatch tracked and fixed by #3363; that mismatch was not the cause of the supervisor shutdown failure.Impact / Why This Matters
Required Kubernetes E2E lanes failed at the first lifecycle stop, including shared v1alpha1, shared v1beta1, managed workspace, workspace operator, and external compute-driver modes. The failures blocked unrelated changes in the merge queue.
Disabling or quarantining the lifecycle scenario would leave Kubernetes stop, start, workspace preservation, and stopped-sandbox deletion uncertified. Extending the deletion timeout alone would only wait longer for cleanup that can no longer reach its boundary.
Acceptance Criteria
sandbox-lifecycleconformance scenario passes in shared v1alpha1, shared v1beta1, managed workspace, workspace operator, and external compute-driver lanes.Reproduction Steps
On an affected revision before #3424:
sandbox-lifecycleconformance scenario.sandbox stopsuspend the workload before supervisor deletion completes.timed out waiting for supervisor Pod deletion.Environment
mainrevisions beginning at8de26878f9324a822131ff6107861134843f1886Validation
PR #3424 implements the corrected sequencing. Its required Kubernetes E2E run passed all affected modes:
The overall workflow was red only because of an unrelated Docker provider-readiness fixture failure inherited from
main; every Kubernetes lane passed.Logs
First reproducing merge-group run: https://github.com/NVIDIA/OpenShell/actions/runs/35244966017
Repeated example: https://github.com/NVIDIA/OpenShell/actions/runs/35251051224/job/105306164605