Skip to content

bug(kubernetes): stop supervisor before suspending workload #3421

Description

@pimlock

User Story

As an OpenShell operator using the Kubernetes compute driver, I want sandbox stop and start operations to complete reliably, so that lifecycle operations work across supported Agent Sandbox API versions and workspace modes.

Problem Statement

Kubernetes sandbox stop used the wrong teardown order for the separately scheduled supervisor introduced by RFC 0012:

  1. The driver suspended the Agent Sandbox workload.
  2. Suspending the workload removed the workload Pod and its sandbox boundary endpoint.
  3. The driver then deleted the supervisor Pod.
  4. The supervisor received SIGTERM and attempted normal backend cleanup through the boundary that had already disappeared.
  5. Cleanup could not complete, so Kubernetes eventually force-killed the supervisor and the driver reported timed out waiting for supervisor Pod deletion.

This is a deterministic sequencing defect, not a timing race between the driver timeout and the Pod termination grace period. The matching 30-second values made the failure present as a deletion-timeout race, but increasing either timeout alone would not restore the boundary required for normal supervisor shutdown.

The correct order is to record durable stop intent, delete the supervisor while the workload boundary remains reachable, wait for supervisor deletion, and only then suspend the workload. The durable phase is required so reconciliation can resume either step after interruption without reversing the order.

The failure is exercised by the sandbox-lifecycle conformance scenario added in #3375. Generation Secret cleanup also encountered a separate Helm RBAC mismatch tracked and fixed by #3363; that mismatch was not the cause of the supervisor shutdown failure.

Impact / Why This Matters

Required Kubernetes E2E lanes failed at the first lifecycle stop, including shared v1alpha1, shared v1beta1, managed workspace, workspace operator, and external compute-driver modes. The failures blocked unrelated changes in the merge queue.

Disabling or quarantining the lifecycle scenario would leave Kubernetes stop, start, workspace preservation, and stopped-sandbox deletion uncertified. Extending the deletion timeout alone would only wait longer for cleanup that can no longer reach its boundary.

Acceptance Criteria

  • Supervisor shutdown begins while the workload boundary remains reachable.
  • Workload suspension begins only after the supervisor Pod is absent.
  • Teardown ordering is represented by durable, idempotent phases that reconciliation can resume after interruption.
  • Supervisor Pod deletion allows its configured termination grace period plus Kubernetes API observation headroom.
  • Generation Secret cleanup remains separate from supervisor shutdown and workload suspension.
  • Regression coverage exercises the ordering, phase transitions, restart/resume behavior, and supervisor termination grace contract.
  • The sandbox-lifecycle conformance scenario passes in shared v1alpha1, shared v1beta1, managed workspace, workspace operator, and external compute-driver lanes.
  • Required Kubernetes E2E continues to run all registered conformance scenarios without a lifecycle quarantine.

Reproduction Steps

On an affected revision before #3424:

  1. Deploy the gateway with the Kubernetes compute driver to the Kubernetes E2E kind cluster.
  2. Run the standalone sandbox-lifecycle conformance scenario.
  3. Observe the first sandbox stop suspend the workload before supervisor deletion completes.
  4. Observe the supervisor remain in termination until Kubernetes force-kills it, followed by timed out waiting for supervisor Pod deletion.

Environment

  • OpenShell: affected main revisions beginning at 8de26878f9324a822131ff6107861134843f1886
  • OS: GitHub Actions Linux runner with kind
  • Runtime, deployment, or integration: Kubernetes compute driver; Agent Sandbox v1alpha1 v0.4.6 and v1beta1 v0.5.0; shared, managed, operator, and external-driver modes

Validation

PR #3424 implements the corrected sequencing. Its required Kubernetes E2E run passed all affected modes:

The overall workflow was red only because of an unrelated Docker provider-readiness fixture failure inherited from main; every Kubernetes lane passed.

Logs

expected: sandbox '<name>' stop succeeds
actual: exit 1 after 61.9s
Error: code: 'Internal error', message: "stop sandbox failed: timed out waiting for supervisor Pod deletion"

First reproducing merge-group run: https://github.com/NVIDIA/OpenShell/actions/runs/35244966017

Repeated example: https://github.com/NVIDIA/OpenShell/actions/runs/35251051224/job/105306164605

Activity

  1. pimlock commented on Sep 17, 2026

    @pimlock
    CollaboratorAuthor

    🏗️ build-plan

    Implementation Plan

    Issue type: fix
    Complexity: Medium
    Confidence: High — the failure and required ordering are clear; the main risk is keeping teardown idempotent under reconciliation races.

    Summary

    Add a crash-resumable stopping-supervisor phase. Keep the workload boundary reachable while Kubernetes terminates the supervisor, then atomically advance to the existing rollback phase and suspend the workload. Give supervisor deletion its configured Pod grace period plus Kubernetes API observation headroom.

    Scope

    • crates/openshell-driver-kubernetes/src/driver.rs: add the durable stop phase, split supervisor Pod deletion from generation Secret cleanup, resume both teardown phases through reconciliation, and calculate the correct deletion deadline.
    • crates/openshell-driver-kubernetes/src/sandbox_runtime.rs: render an explicit supervisor termination grace period and test the Pod contract.
    • architecture/compute-runtimes.md: document that Kubernetes backend cleanup precedes workload suspension and survives gateway restarts.
    • Temporary quarantine files from PR test(kubernetes): quarantine lifecycle conformance #3423: remove only if that PR merges before this fix.

    Implementation Steps

    1. Render an explicit supervisor Pod termination grace period and calculate deletion time as grace plus KUBE_API_TIMEOUT, retaining the Kubernetes 30-second fallback for existing Pods.
    2. Add a resource-version-guarded stopping-supervisor phase that records stop intent without changing the Agent Sandbox desired running state.
    3. Delete the UID-matched supervisor Pod while its workload boundary remains reachable, without deleting generation Secrets yet.
    4. Advance atomically to rolling-back and Suspended or replicas: 0, then wait for workload and Sandbox stop completion.
    5. Clean generation Secrets and clear lifecycle annotations only after workload teardown completes.
    6. Teach periodic reconciliation and repeated stop requests to resume either phase safely after interruption.
    7. Verify the existing sandbox-lifecycle scenario across shared v1alpha1, shared v1beta1, managed workspace, and external-driver Kubernetes lanes.

    Test Plan

    • Unit tests: explicit/default supervisor grace, grace-plus-headroom deadline, API-version-specific stop-begin and stop-advance patches, phase parsing, idempotent/resumed transitions, and Pod rendering.
    • Integration tests: use an existing mocked Kubernetes API fixture for delayed deletion if one exists; do not introduce a new mock framework solely for this issue.
    • E2E tests: run the existing sandbox-lifecycle scenario in the four affected Kubernetes modes and confirm the supervisor no longer exits 137.

    Risks & Open Questions

    • Stop RPCs and periodic reconciliation can race, so deletion and transitions must retain UID and resource-version guards.
    • A genuinely hung supervisor may use its full grace period; Kubernetes remains the trusted fallback, while the driver waits with API headroom.
    • Secret cleanup must remain separate so the RBAC defect in fix(helm): grant secret cleanup permissions #3363 cannot prevent workload suspension.
    • PR test(kubernetes): quarantine lifecycle conformance #3423 already contains the temporary quarantine. Coordinate whether this fix supersedes it or follows it.

    Documentation Impact

    • Update architecture/compute-runtimes.md.
    • No gateway TOML, Helm values, public API, or user-facing compute-driver configuration changes.
    • No SELinux or AppArmor impact is expected because process identity, /proc, execution, visibility, and security contexts are unchanged.

    Revision 1 — initial plan

  2. added
    state:acceptedA maintainer decided OpenShell should pursue this issue
    on Sep 17, 2026
  3. pimlock commented on Sep 17, 2026

    @pimlock
    CollaboratorAuthor

    🏗️ build-from-issue-agent

    Implemented the durable Kubernetes teardown fix in #3424.

    The driver now records a crash-resumable stopping-supervisor phase, removes the supervisor while the workload boundary is still reachable, then suspends the workload and cleans generation Secrets. Supervisor Pod deletion also waits for the configured termination grace period plus Kubernetes API observation headroom.

    Local verification:

    • mise exec -- cargo test -p openshell-driver-kubernetes --lib — 224 passed
    • mise run test — passed
    • mise run pre-commit — passed

    Kubernetes E2E will run in branch CI because the local host has neither k3d nor a configured Kubernetes context.

  4. changed the title [-]bug(kubernetes): supervisor deletion races the Pod termination grace period[/-] [+]bug(kubernetes): stop supervisor before suspending workload[/+] on Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:sandboxSandbox runtime and isolation workstate:acceptedA maintainer decided OpenShell should pursue this issuetest:e2e-kubernetesRequires Kubernetes end-to-end coverage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions