Skip to content

perf(sdk): stop repeating bundle hashing and teardown state reads - #12974

Merged
cv merged 3 commits into
v1from
perf/faster-lifecycle-operations
Oct 11, 2026
Merged

cv merged 3 commits into
v1from
perf/faster-lifecycle-operations

Conversation

@cv

@cv cv commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Part of #12882.

Failure

Nearly all lifecycle test time is spent inside SDK and CLI operations. I measured representative lifecycle tests against a locally built linux_arm64 bundle by logging every OpenTofu subprocess and every Progress::Completed step. Four tests (fabric_deployment::harness_pi, remote_service::managed_pi_applies_without_generation_and_refused_sandbox_changes_keep_intent, deployment::web_search_cli_export_reapply_and_destroy, deployment::rejected_policy_fails_promptly_with_context_and_allows_recovery_or_destroy) ran 30 SDK or CLI processes and 102.4 s of OpenTofu and bundle work:

Step Total Count Mean
root tofu plan 29.0 s 29 1.0 s
runtime tofu plan 20.9 s 13 1.6 s
root tofu apply 16.4 s 18 0.9 s
export tofu plan -refresh-only 11.8 s 13 0.9 s
bundle.verify 6.8 s 60 114 ms (190–360 ms for each process's first check)
tofu show -json (bindings, saved plans, readiness) 12.9 s 115 110 ms
tofu init 4.7 s 89 50 ms

Two items in that profile are repeated work:

  • Every CLI process hashes the whole bundle (365 MB in 12 files) one file at a time.
  • Destroy and plan-destroy read each stage's bindings with tofu init and tofu show, then read them again before planning that stage.

Decision

  • Hash bundle files on several threads. The first failure in manifest order still decides the error.
  • Plan each teardown stage from the bindings already read. Nothing changes that stage's state between the two reads.

Both changes leave results and errors the same.

The largest remaining fixed cost is outside this change. The pinned kreuzwerker/docker 4.6.0 provider waits a fixed 2 s on every docker_network read and removal (networkReadRefreshDelay and networkRemoveRefreshDelay; upstream master still has the same constants). Runtime plans, refreshes, creating applies, and destroys of a Docker-managed service each pay it. In remote_service::remote_model_reapply_and_drift_keep_bindings_and_stop_on_observation_failure, 11 waits account for about 22 s of 62 s. Removing it needs a decision on how the network is managed:

  • Manage the network with a NemoClaw provider resource, and move existing docker_network state to it without recreating the network.
  • Ask upstream to drop the fixed delay, then bump the pinned version.
  • Build a patched Docker provider into the bundle.

Validation

  • New bundle::tests::every_listed_file_is_verified_and_the_first_listed_failure_is_reported pins verification of every listed file and the error order. It passed before and after the change.

  • deployment::destroy_does_not_require_the_inference_credential_or_rewrite_its_reference now also asserts that plan-destroy and destroy each initialize the stage twice: once to read bindings and once for the teardown graph. It failed before the change (3 initializations) and passes after.

  • Locally: the SDK bundle:: and deployment:: unit tests; lifecycle tests for teardown, Helm recovery, and the remote_service::managed_bearer_* scenarios; cargo fmt --all --check; and cargo clippy --workspace --all-targets -- -D warnings.

  • The same four tests after the change: 96.3 s of OpenTofu and bundle work (was 102.4 s); bundle.verify 1.8 s (was 6.8 s); 49.4 s wall (was 55.7 s).

  • CI lifecycle JUnit reports, run 38086127712, compared with the median of the 9 preceding successful Rust desired-state runs (wall seconds / summed test seconds):

    Partition Median before This PR
    linux_arm64 / 1 206.2 / 821 198.1 / 787
    linux_arm64 / 2 185.4 / 741 180.3 / 720
    linux_amd64 / 1 228.0 / 907 223.2 / 887
    linux_amd64 / 2 204.0 / 814 177.5 / 709

    The linux_arm64 partitions varied by at most 3.7 s across those 9 runs, so their 8 s and 5 s drops are outside the noise. The linux_amd64 partitions varied by up to 31 s, so one run there does not show a change.

cv added 3 commits October 10, 2026 13:37
Every CLI process verifies the bundle before its first operation. The
linux_arm64 bundle lists 365 MB in 12 files, and hashing them one at a
time took 190-360 ms per process. Hash them on several threads instead;
the first failure in manifest order still decides the error.
Destroy and plan-destroy read every stage's bindings with tofu init and
show, then read them again before planning the stage: two more OpenTofu
processes per stage, about 150 ms each. Nothing changes that state in
between, so plan each stage from the bindings already read.
@cv cv added area: performance Latency, throughput, resource use, benchmarks, or scaling v1 NemoClaw v1 branch labels Oct 10, 2026
@github-actions

Copy link
Copy Markdown
Contributor

@cv
cv merged commit e251c18 into v1 Oct 11, 2026
54 checks passed
@cv
cv deleted the perf/faster-lifecycle-operations branch October 11, 2026 00:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: performance Latency, throughput, resource use, benchmarks, or scaling v1 NemoClaw v1 branch

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant