Skip to content

test(sdk): pin Kubernetes recovery behavior under injected failures - #12909

Open
cuongdangn wants to merge 5 commits into
v1from
test-12731-failure-injection-recovery
Open

cuongdangn wants to merge 5 commits into
v1from
test-12731-failure-injection-recovery

Conversation

@cuongdangn

Copy link
Copy Markdown

Adds 40 tests: 30 for how the SDK's Kubernetes operations behave when the cluster API fails or an earlier run was interrupted, 8 for the fault helper that they use, and 2 for document validation and the plan check. It also renames and extends one OpenTofu test. Together they cover scenarios 2, 3, 8, 9, 10 and 11, and the corrupt-receipt half of 4, on the list in #12731; scenario 1 is done in #12825. Scenarios 5, 6, 7 and 12 and the deleted-receipt half of 4 are not in this change. Production code is unchanged; src/deployment/plan.rs gains only a #[cfg(test)] module.

Failure

Applying the storage and authentication resources creates a Namespace, a KEK Secret and a development issuer (six objects) in the user's cluster and records each by UID in a local receipt. docs/design/architecture.md says a failed or incomplete observation stops and preserves the prior binding, and that ownership is checked during observation and again before mutation. #12731 asks to verify that authentication, transport and incomplete-observation failures preserve resource bindings and persistent state, and that ownership is verified before recovery changes resources. The tests group this as three invariants, each with a regression they are written to catch:

  • A failed observation is not confirmed absence. A read error turned into None would pass as absence (the gateway StatefulSet, an issuer object during teardown, the Agent Sandbox CRD), and an answer without a uid would pass as an identity.
  • Ownership is verified before mutation. A re-run after a lost create response could delete, replace or adopt an unrecorded object, and a changed OpenShift UID or group range could let a re-apply write issuer objects.
  • Persistent state is retained. A corrupt receipt could be replaced or deleted, expired development certificates could be regenerated, and a plan that replaces a bound resource could be accepted.

Decision

Most tests call Operations and Development against the fake Kubernetes API that SDK tests already use, wrapped by tests/support/kube_faults.rs. A rule decides per request whether to fail it with an HTTP status or a dropped connection, to apply a write and then drop the connection (a lost response), or to apply it and answer without the object's uid. The wrapper logs every request, and Snapshot compares the cluster objects and the contents of every state file before and after. check_plan and document validation are tested directly. kube_api.rs changes only by making status public. The older fault tests in kubernetes_storage.rs and kubernetes_operations.rs keep their own fixtures; moving them onto kube_faults.rs is a separate change.

The tests are characterization tests: they pass on the unchanged production code, so they pin current behavior and need no production change. No behavior changes, so there is no failing test to observe first; instead, temporary changes to production code show that the tests fail when the behavior they guard breaks (see Validation). Where production is inconsistent, a test pins what it finds and says so: a 500 on the kube-system read and on the StorageClass list is Transport, where the other reads report Query.

Coverage

Paths are under crates/nemoclaw-sdk unless noted. Scenario numbers are those on #12731.

File Scenarios Tests Covers
tests/support/kube_faults.rs helper 8 Fault server, request log, Snapshot, platform fixtures
tests/kubernetes_read_failures.rs 2 16 Platform reads under 401, 403, 500 and a dropped connection (#12825 covers the kube-system and StorageClass reads for 401, 403 and a dropped connection); answers without a uid; a client TLS Secret missing one entry
tests/kubernetes_lost_create_response.rs 3 8 Re-run after a lost create: Namespace, KEK Secret, six issuer objects
tests/kubernetes_corrupt_receipt.rs 4 (corrupt receipt) 2 Truncated JSON and an unknown field, in every operation that reads the receipt
tests/kubernetes_expired_authentication.rs 9 2 Expired development certificates in Development and Operations
tests/kubernetes_openshift_reapply.rs 10 2 Re-apply after authentication is removed, with a changed or missing UID or group range, and an unchanged control
tests/deployment.rs 8 1 An invalid document stops plan and apply before the bundle opens or state is created
src/deployment/plan.rs 11 1 check_plan rejects replacing a bound platform resource
crates/nemoclaw-e2e/tests/provider_protocol.rs 11 0 (1 extended) OpenTofu taints a create that returns state with an error and plans delete plus create

Validation

Run on 0906b1b03d, based on v1 at a46f63be1d, darwin_arm64:

  • cargo fmt --check and cargo clippy --locked --workspace --all-targets -- -D warnings: clean.
  • The 40 new tests (39 in the integration binary, 1 in the library unit tests) under cargo nextest run --locked --profile ci -p nemoclaw-sdk, filtered to them: 40 passed in each of 5 consecutive runs and in 3 runs with all 40 at once (--test-threads 40).
  • NEMOCLAW_TEST_TOFU=<absolute path to dist/darwin_arm64/libexec/tofu> cargo test --locked -p nemoclaw-e2e --test integration provider_protocol::real_tofu -- --ignored: 3 passed, including the extended test. tofu version reports OpenTofu v1.12.6. The lifecycle step of cargo ci runs all of provider_protocol, including these three.
  • Windows approximation: with cfg(unix) forced false in the SDK tests, cargo clippy -p nemoclaw-sdk --tests -- -D warnings is clean and the 11 tests that remain pass. This checks unused items and imports; it is not a Windows build.
  • cargo ci: passed on 0906b1b03d (856 s): fmt, clippy, build, test (1311 passed), schema, bundle and lifecycle (112 passed).

Mutation probes: 25 temporary changes to production code, each reverted so the diff holds none. Each made at least one new test fail, and every new test outside kube_faults.rs failed under at least one of them. They include a read error mapped to None, a 409 on create mapped to Query, a re-run that adopts, or deletes and re-creates, the object that a 409 reports, a 401 mapped to Query, a receipt that does not parse treated as empty, the certificate validity check ignored, expired material regenerated, a changed OpenShift identity accepted, a StatefulSet or cluster identity without a uid adopted, a failed read mapped to Incomplete at each of five reads, connect writing the private key first, validation errors ignored, and check_plan relaxed for every provider resource.

Limits of this evidence

  • The fake API serves plain HTTP on loopback. It does not show RBAC, TLS, token validation, how an API server answers, or the kube client's retries of 429, 503 and 504, which are not injected. A response cut off mid-body is not covered. The OpenShift tests write the range annotations themselves.
  • No faulted read is retried after its fault clears, and no test shows how an operator recovers from a refusal. The controls in the lost-create, OpenShift and expiry tests (unrecorded object removed, ranges unchanged, a certificate window that has not ended) show only that each refusal comes from its stated cause.
  • The answers without a uid are synthetic: an API server assigns every object a uid. They show that the SDK refuses to record or adopt such an answer, not that one occurs. The client TLS Secret that lacks a key is stored in the fake API; whether the chart can leave one out is not tested.
  • Three reads of a path that a call reads more than once (two OpenShift Namespace reads and the second StatefulSet read in teardown) are chosen by the request before them. A change that puts another request in between makes those tests fail with "the fault must fire exactly once".
  • The expiry tests rewrite certificates for the same keys with a window in the past, because the SDK reads the system clock and offers no way to set it. Only the combined expiry of the CA and server certificates is covered; a CA that expires alone, and how the gateway behaves with expired certificates, are not.
  • The OpenTofu test uses the fixture provider, so it shows OpenTofu's behavior only; no test here drives the Kubernetes provider's own partial create. Reading the code, not running it: when a create fails after the receipt records an object, KubernetesBackend::ensure returns that object's identity with the error, OpenTofu taints the instance, and nemoclaw apply stops at the plan check before it reaches the refusal that the lost-create tests pin.
  • Windows was not run locally. Windows CI (Test / windows_amd64) runs 11 of the 40 tests. The other 29 are cfg(unix): 28 in kubernetes_read_failures.rs, kubernetes_lost_create_response.rs, kubernetes_corrupt_receipt.rs and kubernetes_openshift_reapply.rs (#![cfg(unix)], as kubernetes_operations.rs is), and 1 in the operations module of kubernetes_expired_authentication.rs.

Observations for maintainers

  • A lost create response leaves an unrecorded object. A re-run refuses it with BindingMismatch, and no SDK path deletes an unrecorded object (Cluster::delete runs only for recorded issuer objects), so it stays until someone removes it. The tests pin the refusal; a change that adds adoption or cleanup after ownership verification will change them.
  • OpenTofu taints an instance after a create that returns state with an error, and the next plan is delete plus create. The SDK rejects a plan with those actions for bound platform resources (tested with a constructed plan, not OpenTofu output); how a user recovers from it is not tested.
  • The first ensure of storage saves the cluster binding (server address and kube-system uid) before it reads the Agent Sandbox CRD, controller and StorageClass, so a failed read of one leaves a receipt that holds only that binding. The scenario list on [Kubernetes/OpenShift] Broaden v1 recovery tests and document inference limitations #12731 left this unselected; three tests pin it.
  • A 500 on the kube-system read and on the StorageClass list is Transport (read_failure in storage.rs); Cluster::get and the Helm release listing report Query. A recorded gateway whose StatefulSet has no uid is BindingMismatch, where first adoption of the same answer is Incomplete; only the latter is pinned.
  • Development certificates are valid for 365 days and the SDK does not rotate them. After expiry, read and ensure of the authentication resource and connect fail with Authentication; teardown does not load the certificates, and no test covers it.
  • connect writes each credential it decodes to state/client before it reads the next, so a Secret that lacks one leaves the earlier ones. The private key is written last.

Refs #12731

🤖 Generated with Claude Code

Add 40 tests: 30 for how the SDK's Kubernetes operations behave when the
cluster API fails or an earlier run was interrupted, 8 for the fault
helper they use, and 2 for document validation and the plan check. The
failures are HTTP 401, 403 and 500, a dropped connection, a lost create
response, a corrupt receipt, expired development certificates, a changed
OpenShift range, and a plan that replaces a bound resource. A
fault-injecting wrapper around the fake Kubernetes API logs every
request, and Snapshot checks that objects and state files are unchanged.
The tests pass on unchanged production code, so they pin it; 25
temporary production changes each made at least one of them fail. One
OpenTofu test is renamed and also asserts that a create that returns
state with an error is tainted and that the next plan is delete plus
create. Refs #12731.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant