Repository navigation
Conversation
|
🌿 Preview your docs: https://nvidia-preview-pr-3265.docs.buildwithfern.com/openshell |
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
e2b8a35 to
797c36e
Compare
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
|
All contributors have signed the DCO ✍️ ✅ |
ab4bb69 to
9be1774
Compare
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
|
Added the startup policy preparation exchange in Before this change, the supervisor discovered or enriched the startup policy, called the unary The initial
There is no reconnect after preparation. Later reconnects omit the image policy, skip startup preparation, and receive the current gateway bootstrap directly. This keeps the gateway authoritative without requiring it to know what paths exist in the sandbox image. It also removes the old ambiguity around mutation responses: neither the candidate nor the preparation response becomes runtime state. Failure behavior is fail-closed. Candidate mismatches, preparation failures, and invalid prepared policies produce The change includes protocol bindings, architecture documentation with a sequence diagram, positive tests for policy precedence and persistence, a negative invalid-policy test, and the focused Docker live-policy E2E coverage. |
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> # Conflicts: # architecture/gateway.md
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
This comment has been minimized.
This comment has been minimized.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
PR Review StatusThanks @pimlock. I reviewed the current stacked-base merge in critical-only mode against the durable feedback ledger. The merge resolution introduces no new Critical defect, and Blocking findings:
Carried findings:
Gator metadata
|
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
|
Hey @sjenning @mrunalp @derekwaynecarr — this one's been open a bit and is part of the #1731 staged rollout (this is stage 2 of 3). Any chance you have bandwidth for a review pass in the near term? Happy to help unblock if there's anything I can clarify in the meantime. |
Port streamed configuration onto the Stage 1 delivery scheduler and the feature-flag handshake: - Replace protocol revisions with SupervisorHello.supports_config_apply and SessionAccepted.config_apply_enabled. A gateway enables apply only in push mode for supervisors that advertise both snapshot and apply support; other sessions keep polling. Drop the unreleased image_policy hello field. - Drop the Stage 2 FIFO dispatcher and pending queue in favor of the scheduler. Acknowledgement gating holds released updates per component in the registry and delivers them through the session slot. Shadow sessions, which never acknowledge, are not held. - Keep the polling startup path from main, including startup admission reporting and VM identity checks, and load a streamed bootstrap through a separate path that applies the same global-policy enrichment rule. - Port the required-bootstrap timeout, delivery dispositions, and owner reconciler into the config_delivery module. - Keep Stage 1 published docs; the streamed startup flow is internal to the push rollout. Scope the debug skill guidance to push mode. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
- Close the startup session when the gateway does not enable apply and reconnect after the workload starts, as polling supervisors always have. Retry transient connection failures while preparing the startup session. - Mark apply-capable supervisors ready only on SupervisorRuntimeReady, in any delivery mode, and resend it after a reconnect without apply. - Resume polling in a stream-started runtime whenever its current session does not apply configuration, such as after a rollback to poll mode. - Leave restart bookkeeping to runtime readiness instead of admission. - Keep the owner reconciler local to each replica and suppress unchanged snapshots for shadow sessions. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Bring in per-component snapshot build deadlines and the latest main, including consistent-hash supervisor session placement. Session redirect took GatewayMessage field 6 and SupervisorHello fields 5 and 6, so the unreleased streamed-apply fields move after them: startup_config_candidate (8), configuration_admission (9), image_policy_discovery (8), and supports_config_apply (9). The gateway decides a redirect before it builds the required bootstrap or sends a startup policy candidate, so a redirect never consumes the supervisor's one-time policy preparer. The supervisor follows redirects during startup preparation as well as on reconnect, and continues the connection epoch from the prepared session so a reconnect can supersede the session it replaces. The required bootstrap keeps its 45-second deadline; it does not hold a delivery build slot. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
…hots Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Stage 2 lets supervisors apply streamed configuration, so the config enum, field, and --config-delivery-mode help no longer describe push as a shadow-only stream. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Add an opt-in push path for gateway-owned supervisor configuration, selected by the gateway setting config_delivery_mode = "poll" | "push". Poll stays the default, and polling keeps working for older supervisors and gateways. Delivery: - Add bootstrap, snapshot, update, result and admission contracts on ConnectSupervisor. Supervisors advertise supports_config_snapshots and supports_config_apply; the gateway answers with config_apply_enabled. - Configuration delivery is a coalescing DeliveryQueue drained by one on-demand delivery worker. Admission never drops updates, fanouts walk connected sandboxes as cursors, and sandbox-scoped work keeps reserved build slots. - Provider changes rebuild only the sandboxes that attach the provider. - Across replicas, other gateways send the session owner a secret-free PeerNotifyConfigUpdate and the owner rebuilds from shared state. An owner reconciler repairs missed notifications. Streamed apply: - Supervisors that enable apply prepare the startup policy against the image, apply the bootstrap before starting the workload, apply and acknowledge live updates, and stop polling. - A failed live registry reload keeps the last-known-good registry. Durable completion: - Sandbox policy and settings updates commit desired state and a durable ConfigUpdateOperation atomically. Callers choose wait_mode COMMIT_ONLY or WAIT_FOR_COMPLETION, with idempotency keys and GetConfigUpdateOperation. - Operations resolve as applied, degraded, failed, superseded, inactive or cancelled, and survive client timeouts and gateway restarts. This combines the previously stacked changes from #3244, #3265 and #3273. Refs #1731 Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Summary
Initialize supervisors and apply live configuration updates through
ConnectSupervisorwhen the gateway runs withconfig_delivery_mode = "push". Supervisors advertisesupports_config_apply; the gateway answers withconfig_apply_enabled, requires a successful bootstrap, and marks the supervisor ready only after it reports runtime readiness. In poll mode, which stays the default, supervisors keep polling. The supervisor must connect to every remote middleware required by the initial effective policy before starting the workload, even whenfail_openis configured. A failed live registry reload retains the last-known-good registry and keeps the workload running.Stage 2 architecture walkthrough
Related Issue
Part of #1731.
Stacked on #3244.
Changes
supports_config_applyandconfig_apply_enabled, opt-in through push mode.Testing
mise run pre-commit;openshell-server,openshell-supervisor, andopenshell-supervisor-processtests;mise run go:ci;mise run sdk:ts:cimise run ciandmise run e2e:dockerhave not been rerun since the port onto the Stage 1 delivery queue.Known reliability issue observed under load
A focused reproduction with 32 running sandboxes and 32 concurrent exec clients produced 6 false exit-code-1 results across 37,376 exec calls, including commands that explicitly ran
exit 0. Diagnostic logs confirmed that the orphan reaper collected a child's exit status 0 before SSH's waiter receivedECHILDand substituted exit code 1. No policy changes or gateway restart were needed.The implicated spawn/register/reap code predates this PR. Existing fix commit
c5cf4ec2bis included in #3142.Follow-ups
Retry failed builds with backoff, record a durable fanout watermark, reconcile owner handoffs, and roll back builds whose session owner is briefly stale. The 30-second owner reconciler covers missed deliveries until then.
Checklist