User Story
As an operator building an event-driven controller, I want to stop only the sandbox execution that produced an observation, so delayed events cannot stop a newer execution.
Problem Statement
The public StopSandboxRequest accepts workspace_scope, name, and request_id. It cannot express an expected sandbox incarnation or execution. The request ID provides durable at-most-once admission; it is not a target precondition.
The live public Go SDK reproduction demonstrated this scenario against an unpatched Gateway at 374c035 with SQLite, for both stop/start and same-name delete/recreate:
- Execution A produces an observation that is queued for processing.
- The sandbox is stopped and started, producing execution B under the same sandbox identity and name.
- The controller processes A's observation and requests a stop by workspace and name. It cannot require the Gateway to reject the request if B is now current.
Deleting and recreating a sandbox with the same name presents a similar stale-target problem.
Impact / Why This Matters
Controllers must either disable automatic stop or risk stopping a replacement execution. Reading current state before stopping leaves a race between the read and the mutation. Direct Kubernetes Pod UID checks/deletes are backend-specific and bypass the Gateway lifecycle contract.
This blocks portable automation that must bind an action to the execution observed. The finding is a missing API capability; this investigation has not established an authorization bypass or sandbox-isolation vulnerability.
Proposed Design
Retain the public workspace/name addressing convention and support a conditional stop workflow:
- A controller receives trustworthy, Gateway/supervisor-attested execution identity with the relevant extension observation. Define its lifecycle semantics, including stop/start and runtime replacement.
- The controller supplies an expected-target precondition when requesting stop. An opaque token is one possible public representation; internal identifiers need not become public entity references.
- The Gateway checks that precondition as part of the serialized lifecycle operation. A stale or unverifiable target produces a distinguishable result without stopping the current replacement.
- Execution identity changes on explicit start, automatic runtime replacement, and recovery launch attempts that may reconstruct compute. Credential refresh, supervisor reconnection, and replica handoff preserve it when compute is unchanged. Persistent SSH host identity is separate and survives runtime replacement.
- Successful stop receipts and Gateway-confirmed, effect-free stale-target rejections replay their original results for 24 hours when a request ID is supplied. Reusing the request ID with another execution identity is a payload conflict. Driver errors or interrupted operations remain unresolved; they must not be recorded as proven no-effect rejections.
- The initial implementation supports one Gateway process with SQLite. PostgreSQL, including single-replica PostgreSQL, and HA deployments explicitly reject conditional stopping until distributed lifecycle fencing is established. Supervisor-session ownership alone does not provide that fencing.
- Upgrade handling must fail closed for runtime records without an authenticated bootstrap execution binding. Operators establish a fresh binding through a Gateway-controlled stop/start with matching runtime releases and verify new attested observations before enabling automatic responses.
The exact field names, token representation, and internal implementation are open for maintainers to choose.
Acceptance Criteria
Alternatives Considered
- Client-side read/check before stop: cannot close the read-to-mutation race.
- Use only the persistent sandbox ID: distinguishes same-name recreation, but not stop/start of the same sandbox.
- Use backend-native deletion: sacrifices portability and Gateway lifecycle semantics.
- Use a durable stop hold: addresses subsequent activation, but does not establish which execution a delayed stop should target.
Agent Investigation
Original source investigation reviewed main at 912a077bd641272016fb8b2fd58209f6c7c6f194:
Related: #3556 concerns driver-facing sandbox references; #3153 concerns durable activation holds. Neither defines this conditional execution-targeting workflow.
Checklist
Upstream Refresh (2026-10-08)
Reviewed upstream fe3942f2d. The public stop API still has no execution precondition. PR #4235 revalidates delayed internal driver exit observations; it does not fence controller-issued stop requests. PR #4094 introduces persistent SSH identity, while #3661 and #4321 introduce supervisor placement and handoff. PR #4217 must preserve those behaviors alongside execution targeting. The independent policy-version change remains in #4219.
User Story
As an operator building an event-driven controller, I want to stop only the sandbox execution that produced an observation, so delayed events cannot stop a newer execution.
Problem Statement
The public
StopSandboxRequestacceptsworkspace_scope,name, andrequest_id. It cannot express an expected sandbox incarnation or execution. The request ID provides durable at-most-once admission; it is not a target precondition.The live public Go SDK reproduction demonstrated this scenario against an unpatched Gateway at
374c035with SQLite, for both stop/start and same-name delete/recreate:Deleting and recreating a sandbox with the same name presents a similar stale-target problem.
Impact / Why This Matters
Controllers must either disable automatic stop or risk stopping a replacement execution. Reading current state before stopping leaves a race between the read and the mutation. Direct Kubernetes Pod UID checks/deletes are backend-specific and bypass the Gateway lifecycle contract.
This blocks portable automation that must bind an action to the execution observed. The finding is a missing API capability; this investigation has not established an authorization bypass or sandbox-isolation vulnerability.
Proposed Design
Retain the public workspace/name addressing convention and support a conditional stop workflow:
The exact field names, token representation, and internal implementation are open for maintainers to choose.
Acceptance Criteria
Alternatives Considered
Agent Investigation
Original source investigation reviewed
mainat912a077bd641272016fb8b2fd58209f6c7c6f194:Related: #3556 concerns driver-facing sandbox references; #3153 concerns durable activation holds. Neither defines this conditional execution-targeting workflow.
Checklist
Upstream Refresh (2026-10-08)
Reviewed upstream
fe3942f2d. The public stop API still has no execution precondition. PR #4235 revalidates delayed internal driver exit observations; it does not fence controller-issued stop requests. PR #4094 introduces persistent SSH identity, while #3661 and #4321 introduce supervisor placement and handoff. PR #4217 must preserve those behaviors alongside execution targeting. The independent policy-version change remains in #4219.