Skip to content

feat(sandbox): support stop requests conditional on sandbox execution identity #4009

Description

@danehans

User Story

As an operator building an event-driven controller, I want to stop only the sandbox execution that produced an observation, so delayed events cannot stop a newer execution.

Problem Statement

The public StopSandboxRequest accepts workspace_scope, name, and request_id. It cannot express an expected sandbox incarnation or execution. The request ID provides durable at-most-once admission; it is not a target precondition.

The live public Go SDK reproduction demonstrated this scenario against an unpatched Gateway at 374c035 with SQLite, for both stop/start and same-name delete/recreate:

  1. Execution A produces an observation that is queued for processing.
  2. The sandbox is stopped and started, producing execution B under the same sandbox identity and name.
  3. The controller processes A's observation and requests a stop by workspace and name. It cannot require the Gateway to reject the request if B is now current.

Deleting and recreating a sandbox with the same name presents a similar stale-target problem.

Impact / Why This Matters

Controllers must either disable automatic stop or risk stopping a replacement execution. Reading current state before stopping leaves a race between the read and the mutation. Direct Kubernetes Pod UID checks/deletes are backend-specific and bypass the Gateway lifecycle contract.

This blocks portable automation that must bind an action to the execution observed. The finding is a missing API capability; this investigation has not established an authorization bypass or sandbox-isolation vulnerability.

Proposed Design

Retain the public workspace/name addressing convention and support a conditional stop workflow:

  • A controller receives trustworthy, Gateway/supervisor-attested execution identity with the relevant extension observation. Define its lifecycle semantics, including stop/start and runtime replacement.
  • The controller supplies an expected-target precondition when requesting stop. An opaque token is one possible public representation; internal identifiers need not become public entity references.
  • The Gateway checks that precondition as part of the serialized lifecycle operation. A stale or unverifiable target produces a distinguishable result without stopping the current replacement.
  • Execution identity changes on explicit start, automatic runtime replacement, and recovery launch attempts that may reconstruct compute. Credential refresh, supervisor reconnection, and replica handoff preserve it when compute is unchanged. Persistent SSH host identity is separate and survives runtime replacement.
  • Successful stop receipts and Gateway-confirmed, effect-free stale-target rejections replay their original results for 24 hours when a request ID is supplied. Reusing the request ID with another execution identity is a payload conflict. Driver errors or interrupted operations remain unresolved; they must not be recorded as proven no-effect rejections.
  • The initial implementation supports one Gateway process with SQLite. PostgreSQL, including single-replica PostgreSQL, and HA deployments explicitly reject conditional stopping until distributed lifecycle fencing is established. Supervisor-session ownership alone does not provide that fencing.
  • Upgrade handling must fail closed for runtime records without an authenticated bootstrap execution binding. Operators establish a fresh binding through a Gateway-controlled stop/start with matching runtime releases and verify new attested observations before enabling automatic responses.

The exact field names, token representation, and internal implementation are open for maintainers to choose.

Acceptance Criteria

  • Supported extension observations expose authenticated metadata sufficient to bind a stop to the observed sandbox incarnation and execution.
  • A matching conditional request stops the intended execution through the normal authorized Gateway lifecycle.
  • A condition from before stop/start or same-name delete/recreate cannot stop the replacement, including when replacement races with the request.
  • Missing, stale, or unverifiable conditions in the conditional workflow never silently fall back to an unconditional stop.
  • Successful and stale-target retries preserve the original target and terminal result without affecting a replacement; changed payloads cannot reuse a request ID.
  • Uncertain driver failures remain unresolved and cannot impersonate a replayable stale-target rejection.
  • Reconnection and session handoff preserve execution identity; runtime replacement changes execution identity while retaining the sandbox SSH identity.
  • Unsupported persistence deployments reject conditional stop; legacy-runtime upgrade handling is documented and tested.
  • The contract documents identity lifetime, result semantics, and compatibility with existing unconditional stop, and has regression coverage plus at least one real-driver end-to-end test.

Alternatives Considered

  • Client-side read/check before stop: cannot close the read-to-mutation race.
  • Use only the persistent sandbox ID: distinguishes same-name recreation, but not stop/start of the same sandbox.
  • Use backend-native deletion: sacrifices portability and Gateway lifecycle semantics.
  • Use a durable stop hold: addresses subsequent activation, but does not establish which execution a delayed stop should target.

Agent Investigation

Original source investigation reviewed main at 912a077bd641272016fb8b2fd58209f6c7c6f194:

Related: #3556 concerns driver-facing sandbox references; #3153 concerns durable activation holds. Neither defines this conditional execution-targeting workflow.

Checklist

  • I've reviewed existing issues and the published docs.
  • This is a design proposal, not a "please build this" request.

Upstream Refresh (2026-10-08)

Reviewed upstream fe3942f2d. The public stop API still has no execution precondition. PR #4235 revalidates delayed internal driver exit observations; it does not fence controller-issued stop requests. PR #4094 introduces persistent SSH identity, while #3661 and #4321 introduce supervisor placement and handoff. PR #4217 must preserve those behaviors alongside execution targeting. The independent policy-version change remains in #4219.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions