Repository navigation
test(sdk): pin Kubernetes recovery behavior under injected failures - #12909
Open
cuongdangn wants to merge 5 commits into
Open
cuongdangn wants to merge 5 commits into
cuongdangn wants to merge 5 commits into
Conversation
Add 40 tests: 30 for how the SDK's Kubernetes operations behave when the cluster API fails or an earlier run was interrupted, 8 for the fault helper they use, and 2 for document validation and the plan check. The failures are HTTP 401, 403 and 500, a dropped connection, a lost create response, a corrupt receipt, expired development certificates, a changed OpenShift range, and a plan that replaces a bound resource. A fault-injecting wrapper around the fake Kubernetes API logs every request, and Snapshot checks that objects and state files are unchanged. The tests pass on unchanged production code, so they pin it; 25 temporary production changes each made at least one of them fail. One OpenTofu test is renamed and also asserts that a create that returns state with an error is tainted and that the next plan is delete plus create. Refs #12731. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Contributor
|
v1 documentation preview: https://nvidia-preview-nemoclaw-v1-pr-12909.docs.buildwithfern.com/nemoclaw |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds 40 tests: 30 for how the SDK's Kubernetes operations behave when the cluster API fails or an earlier run was interrupted, 8 for the fault helper that they use, and 2 for document validation and the plan check. It also renames and extends one OpenTofu test. Together they cover scenarios 2, 3, 8, 9, 10 and 11, and the corrupt-receipt half of 4, on the list in #12731; scenario 1 is done in #12825. Scenarios 5, 6, 7 and 12 and the deleted-receipt half of 4 are not in this change. Production code is unchanged;
src/deployment/plan.rsgains only a#[cfg(test)]module.Failure
Applying the storage and authentication resources creates a Namespace, a KEK Secret and a development issuer (six objects) in the user's cluster and records each by UID in a local receipt.
docs/design/architecture.mdsays a failed or incomplete observation stops and preserves the prior binding, and that ownership is checked during observation and again before mutation. #12731 asks to verify that authentication, transport and incomplete-observation failures preserve resource bindings and persistent state, and that ownership is verified before recovery changes resources. The tests group this as three invariants, each with a regression they are written to catch:Nonewould pass as absence (the gateway StatefulSet, an issuer object during teardown, the Agent Sandbox CRD), and an answer without auidwould pass as an identity.Decision
Most tests call
OperationsandDevelopmentagainst the fake Kubernetes API that SDK tests already use, wrapped bytests/support/kube_faults.rs. A rule decides per request whether to fail it with an HTTP status or a dropped connection, to apply a write and then drop the connection (a lost response), or to apply it and answer without the object'suid. The wrapper logs every request, andSnapshotcompares the cluster objects and the contents of every state file before and after.check_planand document validation are tested directly.kube_api.rschanges only by makingstatuspublic. The older fault tests inkubernetes_storage.rsandkubernetes_operations.rskeep their own fixtures; moving them ontokube_faults.rsis a separate change.The tests are characterization tests: they pass on the unchanged production code, so they pin current behavior and need no production change. No behavior changes, so there is no failing test to observe first; instead, temporary changes to production code show that the tests fail when the behavior they guard breaks (see Validation). Where production is inconsistent, a test pins what it finds and says so: a 500 on the kube-system read and on the StorageClass list is
Transport, where the other reads reportQuery.Coverage
Paths are under
crates/nemoclaw-sdkunless noted. Scenario numbers are those on #12731.tests/support/kube_faults.rsSnapshot, platform fixturestests/kubernetes_read_failures.rsuid; a client TLS Secret missing one entrytests/kubernetes_lost_create_response.rstests/kubernetes_corrupt_receipt.rstests/kubernetes_expired_authentication.rsDevelopmentandOperationstests/kubernetes_openshift_reapply.rstests/deployment.rsplanandapplybefore the bundle opens or state is createdsrc/deployment/plan.rscheck_planrejects replacing a bound platform resourcecrates/nemoclaw-e2e/tests/provider_protocol.rsValidation
Run on
0906b1b03d, based onv1ata46f63be1d, darwin_arm64:cargo fmt --checkandcargo clippy --locked --workspace --all-targets -- -D warnings: clean.integrationbinary, 1 in the library unit tests) undercargo nextest run --locked --profile ci -p nemoclaw-sdk, filtered to them: 40 passed in each of 5 consecutive runs and in 3 runs with all 40 at once (--test-threads 40).NEMOCLAW_TEST_TOFU=<absolute path to dist/darwin_arm64/libexec/tofu> cargo test --locked -p nemoclaw-e2e --test integration provider_protocol::real_tofu -- --ignored: 3 passed, including the extended test.tofu versionreports OpenTofu v1.12.6. Thelifecyclestep ofcargo ciruns all ofprovider_protocol, including these three.cfg(unix)forced false in the SDK tests,cargo clippy -p nemoclaw-sdk --tests -- -D warningsis clean and the 11 tests that remain pass. This checks unused items and imports; it is not a Windows build.cargo ci: passed on0906b1b03d(856 s): fmt, clippy, build, test (1311 passed), schema, bundle and lifecycle (112 passed).Mutation probes: 25 temporary changes to production code, each reverted so the diff holds none. Each made at least one new test fail, and every new test outside
kube_faults.rsfailed under at least one of them. They include a read error mapped toNone, a 409 on create mapped toQuery, a re-run that adopts, or deletes and re-creates, the object that a 409 reports, a 401 mapped toQuery, a receipt that does not parse treated as empty, the certificate validity check ignored, expired material regenerated, a changed OpenShift identity accepted, a StatefulSet or cluster identity without auidadopted, a failed read mapped toIncompleteat each of five reads,connectwriting the private key first, validation errors ignored, andcheck_planrelaxed for every provider resource.Limits of this evidence
uidare synthetic: an API server assigns every object auid. They show that the SDK refuses to record or adopt such an answer, not that one occurs. The client TLS Secret that lacks a key is stored in the fake API; whether the chart can leave one out is not tested.KubernetesBackend::ensurereturns that object's identity with the error, OpenTofu taints the instance, andnemoclaw applystops at the plan check before it reaches the refusal that the lost-create tests pin.Test / windows_amd64) runs 11 of the 40 tests. The other 29 arecfg(unix): 28 inkubernetes_read_failures.rs,kubernetes_lost_create_response.rs,kubernetes_corrupt_receipt.rsandkubernetes_openshift_reapply.rs(#![cfg(unix)], askubernetes_operations.rsis), and 1 in theoperationsmodule ofkubernetes_expired_authentication.rs.Observations for maintainers
BindingMismatch, and no SDK path deletes an unrecorded object (Cluster::deleteruns only for recorded issuer objects), so it stays until someone removes it. The tests pin the refusal; a change that adds adoption or cleanup after ownership verification will change them.ensureof storage saves the cluster binding (server address and kube-systemuid) before it reads the Agent Sandbox CRD, controller and StorageClass, so a failed read of one leaves a receipt that holds only that binding. The scenario list on [Kubernetes/OpenShift] Broaden v1 recovery tests and document inference limitations #12731 left this unselected; three tests pin it.Transport(read_failureinstorage.rs);Cluster::getand the Helm release listing reportQuery. A recorded gateway whose StatefulSet has nouidisBindingMismatch, where first adoption of the same answer isIncomplete; only the latter is pinned.readandensureof the authentication resource andconnectfail withAuthentication; teardown does not load the certificates, and no test covers it.connectwrites each credential it decodes tostate/clientbefore it reads the next, so a Secret that lacks one leaves the earlier ones. The private key is written last.Refs #12731
🤖 Generated with Claude Code