Repository navigation
feat(kubernetes): request resources for the supervisor container - #4234
Closed
ozbarshalom wants to merge 1 commit into
Closed
ozbarshalom wants to merge 1 commit into
ozbarshalom wants to merge 1 commit into
Conversation
The supervisor container was created without resources, so every supervisor pod ran as BestEffort: first to be evicted under node memory pressure, invisible to namespace ResourceQuota and scheduling, and rejected in namespaces whose quota covers compute requests. Add sandbox_runtime.supervisor_resources to the Kubernetes driver config with default requests of 50m CPU and 64Mi memory and no limits, expose it as supervisor.resources in the Helm chart, and document it. Closes NVIDIA#3415 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Oz Bar Shalom <ozb@nvidia.com>
ozbarshalom
requested review from
a team,
derekwaynecarr,
mrunalp and
sjenning
as code owners
October 6, 2026 11:56
|
Thank you for your submission! We ask that you sign our Developer Certificate of Origin before we can accept your contribution. You can sign the DCO by adding a comment below using this text: I have read the DCO document and I hereby sign the DCO. You can retrigger this bot by commenting recheck in this Pull Request. Posted by the DCO Assistant Lite bot. |
|
Thank you for your interest in contributing to OpenShell, @ozbarshalom. This project uses a vouch system for first-time contributors. Before submitting a pull request, you need to be vouched by a maintainer. To get vouched:
See CONTRIBUTING.md for details. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The Kubernetes supervisor container is created with no
resources, so every supervisor pod isBestEffort. That makes it the first eviction candidate under node memory pressure. It also leaves it uncounted by namespaceResourceQuotaand the scheduler, and gets it rejected in namespaces whose quota covers compute requests. This PR gives the supervisor container default CPU and memory requests (50m/64Mi) and makes its requests and limits configurable in the gateway configuration and Helm chart.Related Issue
Closes #3415
Refs #3930. This covers the supervisor container; the workload pod's init containers are out of scope here.
Changes
[openshell.drivers.kubernetes.sandbox_runtime.supervisor_resources]withrequestsandlimitsmaps (KubernetesContainerResources).cpu = "50m",memory = "64Mi", no limits.validate()rejects empty names or quantities. The API server validates the quantity format when it creates the pod, as it already does for workload resources.supervisor_pod()sets the container'sresourcesfrom that config, and leaves the field unset when both maps are empty.boundary_portand uses the default supervisor resources.supervisor.resources(defaults matching the driver), rendered intogateway.toml. Chart README regenerated withhelm-docs.docs/how-it-works/gateways/configuration.mdxdocuments the setting, its defaults and its QoS and quota effect.The defaults come from measuring a supervisor on a live cluster: about 37m CPU and 24Mi working set while idle and while relaying a full coding-agent session with TLS interception. That leaves headroom without reserving much per sandbox. No default limit is set, to avoid throttling or out-of-memory kills of the component holding the gateway session; operators can add limits where their cluster requires them.
Testing
cargo test -p openshell-driver-kubernetes --lib: 287 passed, including 7 new tests:resourceswhen unconfigured.cargo fmt --checkandcargo clippy -p openshell-driver-kubernetes -p openshell-gateway --all-targets -- -D warningspass.helm unittest deploy/helm/openshell: 260 passed, including 2 new cases for the default and configured rendering.test:e2e-kubernetesrun would confirm the supervisor pod's QoS class on a cluster.Checklist