Skip to content

docs: add removal of a deployment's retained Docker resources - #12929

Open
bowenzhu21 wants to merge 10 commits into
NVIDIA:v1from
bowenzhu21:docs/retained-resource-removal
Open

bowenzhu21 wants to merge 10 commits into
NVIDIA:v1from
bowenzhu21:docs/retained-resource-removal

Conversation

@bowenzhu21

@bowenzhu21 bowenzhu21 commented Oct 9, 2026 •

Copy link
Copy Markdown

Part of #12640.

Failure

Destroy keeps gateway storage, model downloads, and service credentials by design, but the guide had no way to remove them and pointed at #12640 instead. cargo ci live-docker discarded every cleanup error, so anything its tests left behind went unseen.

Decision

Document removal on Docker engines. Remove Retained Resources in usage.md has the reader confirm destroy finished, list the deployment's objects by its UID label on each engine, match them to the entries plan --destroy prints under Retained resources: using a new table in state.md, and remove them, containers first. The state directory stays until every entry's objects are gone, so a missed engine does not look finished. state.md also gains the missing managed Kubernetes row, and pages that mention retained data link the procedure or state what stays open.

cargo ci live-docker now fails, naming the objects, when OpenShell leaves a sandbox object or when a labelled object cannot be listed or survives removal.

A purge command, Podman, Kubernetes, an external gateway's workspace, and verified removal of service volumes and of objects on SSH engines stay under #12640.

Validation

  • A managed deployment applied through the CLI with a sandbox retained OpenShell workspace and gateway storage after destroy, the list brev.rs asserts for a full deployment; its agent did not become ready in this environment. On Linux ARM64 with Docker 29.5.2 in a Colima VM, the section's commands, run as written, listed exactly the initializer, volume, and bridge, removed them, and a rerun changed nothing.
  • The SDK's managed gateway recovery test, which applies only the runtime stage, retained gateway storage alone and was removed the same way. Applying with its old state directory then failed with the recreation guard, and a fresh UID and state directory deployed again.
  • During cargo ci live-docker, OpenShell labelled the sandbox containers and volumes it created with openshell.ai/sandbox-namespace and their gateway's name, which is what the check selects.
  • Unit tests use a fake Docker CLI that needs --all and --force and refuses to remove a volume or network while any container remains. Each of eleven deliberate breaks, such as removing volumes before containers or dropping --all, fails at least one, and 400 repeated runs passed.
  • With v1 at 29e4506 merged, cargo ci passed on Linux ARM64 (1,334 tests, 129 lifecycle tests), and cargo ci live-docker passed all 10 tests and left no labelled or sandbox object. A fork run of CI / Native passed on Linux ARM64 and AMD64 with both live suites, macOS, and Windows, and CI / Images passed on both architectures.

Signed-off-by: Bowen Zhu bowenzhu66@gmail.com

After its tests pass, cargo ci live-docker fails if OpenShell left a sandbox object in one of its gateways' namespaces, or if a container, volume, or network labelled with one of its run UUIDs survives removal. It removes those objects, containers first, whether or not the tests pass, and removes the images it built last.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
Add the removal procedure to the destroy guide, the managed Kubernetes gateway and the Docker object names to the retention reference, and links or limits on the pages that mention retained data.

Verified on Linux ARM64 with Docker 29.5.2 in a Colima VM at 9e2f241, using the SDK's managed gateway recovery test, which applies and destroys only the runtime stage, with no sandboxes or services. After destroy, the initializer, gateway volume, and bridge remained, labelled with the UID and named after its workspace. The section's commands, run as written, found no sandbox objects, removed the three objects, and left the lists empty. A rerun changed nothing, applying with the old state directory failed with the recreation guard, and a new state directory applied again.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 9, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Follow the existing procedures: prose prerequisites, one lead-in per command block, deployment-prefixed variables, the documented UID-derived workspace placeholder instead of a hash command, and the model runtime guide's instruction to run on an SSH-placed host. The section shrinks from four command blocks to two.

Run as written on a destroyed managed Docker gateway (Linux ARM64, Docker 29.5.2 in a Colima VM): the listing showed only the initializer, gateway volume, and bridge, named after the workspace; the sandbox commands printed nothing; removal and a rerun succeeded; and every list command then printed nothing.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
List and remove a deployment's objects by its UID label alone, without the workspace placeholder or the sandbox namespace commands, which the live Docker suite still checks. Name every engine a deployment can select, including a service's SSH placement.engine, and point at Docker's in-use error, which names the container, when a volume or network stays.

Run as written on a destroyed managed Docker gateway (Linux ARM64, Docker 29.5.2 in a Colima VM): the listing showed only the initializer, gateway volume, and bridge, removal and a rerun succeeded, and every list command then printed nothing.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
Writing the fake Docker CLI and executing it from parallel tests raced their forks: on Linux, 33 of 150 runs failed with Text file busy, which the listing reported as a cleanup failure. The fake now runs through sh, and 300 runs passed. It also needs --all for stopped containers, --force for running ones, and the matching kind on rm, so dropping any of them fails a test, and the namespace test compares against the SDK's workspace instead of a literal.

A run whose tests fail now prints any cleanup failure instead of discarding it, listing failures are reported once as cannot list, and Docker's removal errors are no longer hidden.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
Key the Docker object table by the entries plan --destroy prints under Retained resources, and keep the state directory until every entry's objects are gone, so a missed engine no longer looks finished. Say why the UID must be unique and why an old state directory cannot be reused, link the live suite without overclaiming, and drop the either/or from bundle removal.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
@bowenzhu21
bowenzhu21 marked this pull request as ready for review October 9, 2026 20:41
Bundle removal keeps state for anything the removal procedure does not cover, so link the issue that tracks it, as the other pages that state a limit do.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
Bind the sandbox and owner filters by name instead of indexing an array, build the full workspace name as the SDK does, and name the test after what it checks.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
A full managed deployment also retains its OpenShell workspace, which lives in the gateway storage volume, so the table maps that entry and the procedure can finish. Split the closing checks into separate steps, set the engine directly, give each limit its own row, and restore the model diagnostics sentence.

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>
…removal

Signed-off-by: Bowen Zhu <bowenzhu66@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant