Skip to content

feat(kubernetes): request resources for the supervisor container - #4234

Closed
ozbarshalom wants to merge 1 commit into
NVIDIA:mainfrom
ozbarshalom:feat/3415-supervisor-resources/ozbarshalom
Closed

ozbarshalom wants to merge 1 commit into
NVIDIA:mainfrom
ozbarshalom:feat/3415-supervisor-resources/ozbarshalom

Conversation

@ozbarshalom

Copy link
Copy Markdown

Summary

The Kubernetes supervisor container is created with no resources, so every supervisor pod is BestEffort. That makes it the first eviction candidate under node memory pressure. It also leaves it uncounted by namespace ResourceQuota and the scheduler, and gets it rejected in namespaces whose quota covers compute requests. This PR gives the supervisor container default CPU and memory requests (50m / 64Mi) and makes its requests and limits configurable in the gateway configuration and Helm chart.

Related Issue

Closes #3415

Refs #3930. This covers the supervisor container; the workload pod's init containers are out of scope here.

Changes

  • Driver config: [openshell.drivers.kubernetes.sandbox_runtime.supervisor_resources] with requests and limits maps (KubernetesContainerResources).
    • Default: requests cpu = "50m", memory = "64Mi", no limits.
    • Overrides: setting a map replaces the default for that map; an empty table omits it.
    • Validation: validate() rejects empty names or quantities. The API server validates the quantity format when it creates the pod, as it already does for workload resources.
  • Supervisor pod: supervisor_pod() sets the container's resources from that config, and leaves the field unset when both maps are empty.
  • Standalone driver binary: keeps its explicit boundary_port and uses the default supervisor resources.
  • Helm chart: supervisor.resources (defaults matching the driver), rendered into gateway.toml. Chart README regenerated with helm-docs.
  • Docs: docs/how-it-works/gateways/configuration.mdx documents the setting, its defaults and its QoS and quota effect.

The defaults come from measuring a supervisor on a live cluster: about 37m CPU and 24Mi working set while idle and while relaying a full coding-agent session with TLS interception. That leaves headroom without reserving much per sandbox. No default limit is set, to avoid throttling or out-of-memory kills of the component holding the gateway session; operators can add limits where their cluster requires them.

Testing

  • Checks appropriate to the affected code and behavior pass
    • cargo test -p openshell-driver-kubernetes --lib: 287 passed, including 7 new tests:
      • config: defaults; per-map replacement; clearing with an empty table; unknown fields and empty quantities rejected;
      • pod rendering: default requests; configured requests and limits; no resources when unconfigured.
    • The existing test that parses the documented Kubernetes TOML example also covers the new snippet.
    • cargo fmt --check and cargo clippy -p openshell-driver-kubernetes -p openshell-gateway --all-targets -- -D warnings pass.
    • helm unittest deploy/helm/openshell: 260 passed, including 2 new cases for the default and configured rendering.
  • Unit tests added/updated (if applicable)
  • E2E tests added/updated (if applicable): not run locally; a test:e2e-kubernetes run would confirm the supervisor pod's QoS class on a cluster.

Checklist

  • Follows Conventional Commits
  • Commits are signed off (DCO)
  • Architecture docs updated (if applicable): n/a; the gateway configuration reference is updated

The supervisor container was created without resources, so every
supervisor pod ran as BestEffort: first to be evicted under node memory
pressure, invisible to namespace ResourceQuota and scheduling, and
rejected in namespaces whose quota covers compute requests.

Add sandbox_runtime.supervisor_resources to the Kubernetes driver
config with default requests of 50m CPU and 64Mi memory and no limits,
expose it as supervisor.resources in the Helm chart, and document it.

Closes NVIDIA#3415

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Oz Bar Shalom <ozb@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

Thank you for your submission! We ask that you sign our Developer Certificate of Origin before we can accept your contribution. You can sign the DCO by adding a comment below using this text:


I have read the DCO document and I hereby sign the DCO.


You can retrigger this bot by commenting recheck in this Pull Request. Posted by the DCO Assistant Lite bot.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

Thank you for your interest in contributing to OpenShell, @ozbarshalom.

This project uses a vouch system for first-time contributors. Before submitting a pull request, you need to be vouched by a maintainer.

To get vouched:

  1. Open a Vouch Request discussion.
  2. Describe what you want to change and why.
  3. Write in your own words — do not have an AI generate the request.
  4. A maintainer will comment /vouch if approved.
  5. Once vouched, open a new PR (preferred) or reopen this one after a few minutes.

See CONTRIBUTING.md for details.

@github-actions github-actions Bot closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: Kubernetes supervisor Pod is BestEffort QoS, so eviction terminates the sandbox

1 participant