Skip to content

feat(supervisor)!: apply streamed configuration snapshots - #3265

Closed
pimlock wants to merge 41 commits into
1731-config-update-stage-1/pimlockfrom
1731-config-update-stage-2/pimlock
Closed

pimlock wants to merge 41 commits into
1731-config-update-stage-1/pimlockfrom
1731-config-update-stage-2/pimlock

Conversation

@pimlock

@pimlock pimlock commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Initialize supervisors and apply live configuration updates through ConnectSupervisor when the gateway runs with config_delivery_mode = "push". Supervisors advertise supports_config_apply; the gateway answers with config_apply_enabled, requires a successful bootstrap, and marks the supervisor ready only after it reports runtime readiness. In poll mode, which stays the default, supervisors keep polling. The supervisor must connect to every remote middleware required by the initial effective policy before starting the workload, even when fail_open is configured. A failed live registry reload retains the last-known-good registry and keeps the workload running.

Stage 2 architecture walkthrough

Related Issue

Part of #1731.

Stacked on #3244.

Changes

  • Negotiate streamed apply with supports_config_apply and config_apply_enabled, opt-in through push mode.
  • Make the stream authoritative for bootstrap and live configuration, including startup policy preparation, for sessions that enable apply.
  • Deliver through the Stage 1 delivery queue, holding each component until the supervisor acknowledges the previous update.
  • Resume polling when a session does not enable apply, for example after a rollback to poll mode.
  • Record sanitized application outcomes and preserve policy fallback behavior.

Testing

  • mise run pre-commit; openshell-server, openshell-supervisor, and openshell-supervisor-process tests; mise run go:ci; mise run sdk:ts:ci
  • mise run ci and mise run e2e:docker have not been rerun since the port onto the Stage 1 delivery queue.

Known reliability issue observed under load

A focused reproduction with 32 running sandboxes and 32 concurrent exec clients produced 6 false exit-code-1 results across 37,376 exec calls, including commands that explicitly ran exit 0. Diagnostic logs confirmed that the orphan reaper collected a child's exit status 0 before SSH's waiter received ECHILD and substituted exit code 1. No policy changes or gateway restart were needed.

The implicated spawn/register/reap code predates this PR. Existing fix commit c5cf4ec2b is included in #3142.

Follow-ups

Retry failed builds with backoff, record a durable fanout watermark, reconcile owner handoffs, and roll back builds whose session owner is briefly stale. The 30-second owner reconciler covers missed deliveries until then.

Checklist

  • Conventional commits with DCO sign-off
  • Published docs unchanged: streamed apply is internal to the push rollout
  • Related public skill reviewed and updated
  • No generated TypeScript sources committed

@github-actions

Copy link
Copy Markdown

@pimlock
pimlock added this pull request to stack #3266 September 10, 2026 23:38
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@pimlock
pimlock force-pushed the 1731-config-update-stage-2/pimlock branch from e2b8a35 to 797c36e Compare September 11, 2026 02:08
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Sep 16, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

All contributors have signed the DCO ✍️ ✅
Posted by the DCO Assistant Lite bot.

@pimlock
pimlock force-pushed the 1731-config-update-stage-2/pimlock branch from ab4bb69 to 9be1774 Compare September 16, 2026 18:27
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@pimlock

pimlock commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator Author

Added the startup policy preparation exchange in fa58bddb8.

Before this change, the supervisor discovered or enriched the startup policy, called the unary UpdateConfig RPC, then reconnected so it could receive a fresh authoritative bootstrap. That extra mutation and reconnect made the startup ownership model harder to follow.

The initial ConnectSupervisor stream now carries the complete exchange:

  1. SupervisorHello may offer the policy found in the sandbox image.
  2. The gateway selects its existing policy when present. Otherwise it selects the image policy.
  3. The gateway sends the selected full policy as StartupConfigCandidate.
  4. The supervisor performs image-specific filesystem enrichment and replies with StartupConfigPrepared, using unchanged, a complete prepared policy, or a bounded failure.
  5. The gateway treats a prepared policy as an untrusted proposal. It runs it through the normal sandbox policy update path, including validation, safety checks, persistence, and composition.
  6. The gateway rebuilds the bootstrap and sends SessionAccepted. The supervisor initializes only from that bootstrap.

There is no reconnect after preparation. Later reconnects omit the image policy, skip startup preparation, and receive the current gateway bootstrap directly.

This keeps the gateway authoritative without requiring it to know what paths exist in the sandbox image. It also removes the old ambiguity around mutation responses: neither the candidate nor the preparation response becomes runtime state. SessionAccepted remains the single startup source of truth.

Failure behavior is fail-closed. Candidate mismatches, preparation failures, and invalid prepared policies produce SessionRejected before the supervisor can observe SessionAccepted. Session traffic is buffered until the remaining fallible acceptance work completes, so a config update cannot overtake acceptance.

The change includes protocol bindings, architecture documentation with a sequence diagram, positive tests for policy precedence and persistence, a negative invalid-policy test, and the focused Docker live-policy E2E coverage.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@pimlock pimlock added gator:approval-needed Gator completed review; maintainer approval needed and removed gator:blocked Gator is blocked by process or repository gates labels Sep 21, 2026
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

# Conflicts:
#	architecture/gateway.md
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@pimlock pimlock changed the title feat(supervisor): apply streamed configuration snapshots feat(supervisor)!: apply streamed configuration snapshots Sep 21, 2026
@pimlock

This comment has been minimized.

@pimlock pimlock added gator:watch-pipeline Gator is monitoring PR CI/CD status gator:blocked Gator is blocked by process or repository gates and removed gator:approval-needed Gator completed review; maintainer approval needed gator:watch-pipeline Gator is monitoring PR CI/CD status labels Sep 21, 2026
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@pimlock

pimlock commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

gator-agent

PR Review Status

Thanks @pimlock. I reviewed the current stacked-base merge in critical-only mode against the durable feedback ledger. The merge resolution introduces no new Critical defect, and GATOR-c10f3724-01 through GATOR-c10f3724-04 remain resolved.

Blocking findings:

  • No blocking findings remain.

Carried findings:

  • None; all four prior Gator findings remain resolved.
Gator metadata
  • Validation: Project-valid stage 2 of accepted issue Push gateway-owned desired state to supervisors #1731, stacked on active stage-1 PR feat(supervisor): stage gateway configuration snapshot delivery #3244.
  • Docs: The reviewed merge resolution does not introduce a new direct UX change; the PR's existing operator documentation remains present.
  • Checks: Current-head Branch Checks and E2E are queued or running; completed required gates are green.
  • E2E: test:e2e and test:e2e-kubernetes remain applied, and current-head workflows were dispatched without a rerun or /ok to test.
  • Head SHA: 1dc9c881ea2b5ac75a93e6dca5e6ee65e1d0c165
  • Base SHA: 2091c38a87979c98db51a73722e4e2a204906575
  • Merge base SHA: 2091c38a87979c98db51a73722e4e2a204906575
  • Patch ID: 8246581dd660c77df41e64fc0013fc444c069255
  • Gator payload: 9
  • Review mode: critical_only
  • Previous reviewed SHA: bd0ad5ee590142391b8eda9cdbf9ed59ee772638
  • Review budget exhausted: yes
  • Maintainer decision required: no
  • Next state: gator:watch-pipeline

@pimlock pimlock added gator:watch-pipeline Gator is monitoring PR CI/CD status gator:blocked Gator is blocked by process or repository gates and removed gator:blocked Gator is blocked by process or repository gates gator:watch-pipeline Gator is monitoring PR CI/CD status labels Sep 21, 2026
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@politerealism

Copy link
Copy Markdown
Contributor

Hey @sjenning @mrunalp @derekwaynecarr — this one's been open a bit and is part of the #1731 staged rollout (this is stage 2 of 3). Any chance you have bandwidth for a review pass in the near term? Happy to help unblock if there's anything I can clarify in the meantime.

Port streamed configuration onto the Stage 1 delivery scheduler and the
feature-flag handshake:

- Replace protocol revisions with SupervisorHello.supports_config_apply and
  SessionAccepted.config_apply_enabled. A gateway enables apply only in push
  mode for supervisors that advertise both snapshot and apply support; other
  sessions keep polling. Drop the unreleased image_policy hello field.
- Drop the Stage 2 FIFO dispatcher and pending queue in favor of the
  scheduler. Acknowledgement gating holds released updates per component in
  the registry and delivers them through the session slot. Shadow sessions,
  which never acknowledge, are not held.
- Keep the polling startup path from main, including startup admission
  reporting and VM identity checks, and load a streamed bootstrap through a
  separate path that applies the same global-policy enrichment rule.
- Port the required-bootstrap timeout, delivery dispositions, and owner
  reconciler into the config_delivery module.
- Keep Stage 1 published docs; the streamed startup flow is internal to the
  push rollout. Scope the debug skill guidance to push mode.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
- Close the startup session when the gateway does not enable apply and
  reconnect after the workload starts, as polling supervisors always have.
  Retry transient connection failures while preparing the startup session.
- Mark apply-capable supervisors ready only on SupervisorRuntimeReady, in
  any delivery mode, and resend it after a reconnect without apply.
- Resume polling in a stream-started runtime whenever its current session
  does not apply configuration, such as after a rollback to poll mode.
- Leave restart bookkeeping to runtime readiness instead of admission.
- Keep the owner reconciler local to each replica and suppress unchanged
  snapshots for shadow sessions.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Bring in per-component snapshot build deadlines and the latest main,
including consistent-hash supervisor session placement.

Session redirect took GatewayMessage field 6 and SupervisorHello fields
5 and 6, so the unreleased streamed-apply fields move after them:
startup_config_candidate (8), configuration_admission (9),
image_policy_discovery (8), and supports_config_apply (9).

The gateway decides a redirect before it builds the required bootstrap
or sends a startup policy candidate, so a redirect never consumes the
supervisor's one-time policy preparer. The supervisor follows redirects
during startup preparation as well as on reconnect, and continues the
connection epoch from the prepared session so a reconnect can supersede
the session it replaces. The required bootstrap keeps its 45-second
deadline; it does not hold a delivery build slot.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
…hots

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Stage 2 lets supervisors apply streamed configuration, so the config enum, field, and --config-delivery-mode help no longer describe push as a shadow-only stream.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@pimlock

pimlock commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded by #4320, which combines #3244, #3265 and #3273 into one PR so the work can iterate in one place. Review history stays here.

@pimlock pimlock closed this Oct 7, 2026
pimlock added a commit that referenced this pull request Oct 7, 2026
Add an opt-in push path for gateway-owned supervisor configuration,
selected by the gateway setting config_delivery_mode = "poll" | "push".
Poll stays the default, and polling keeps working for older supervisors
and gateways.

Delivery:
- Add bootstrap, snapshot, update, result and admission contracts on
  ConnectSupervisor. Supervisors advertise supports_config_snapshots and
  supports_config_apply; the gateway answers with config_apply_enabled.
- Configuration delivery is a coalescing DeliveryQueue drained by one
  on-demand delivery worker. Admission never drops updates, fanouts walk
  connected sandboxes as cursors, and sandbox-scoped work keeps reserved
  build slots.
- Provider changes rebuild only the sandboxes that attach the provider.
- Across replicas, other gateways send the session owner a secret-free
  PeerNotifyConfigUpdate and the owner rebuilds from shared state. An
  owner reconciler repairs missed notifications.

Streamed apply:
- Supervisors that enable apply prepare the startup policy against the
  image, apply the bootstrap before starting the workload, apply and
  acknowledge live updates, and stop polling.
- A failed live registry reload keeps the last-known-good registry.

Durable completion:
- Sandbox policy and settings updates commit desired state and a durable
  ConfigUpdateOperation atomically. Callers choose wait_mode COMMIT_ONLY
  or WAIT_FOR_COMPLETION, with idempotency keys and GetConfigUpdateOperation.
- Operations resolve as applied, degraded, failed, superseded, inactive or
  cancelled, and survive client timeouts and gateway restarts.

This combines the previously stacked changes from #3244, #3265 and #3273.

Refs #1731

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

gator:blocked Gator is blocked by process or repository gates test:e2e Requires end-to-end coverage test:e2e-kubernetes Requires Kubernetes end-to-end coverage

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants