Repository navigation
Make the patrol oracle fly the lead's autopilot, and floor the slot at the GPS resolution - #60
Merged
Merged
Conversation
…t the GPS resolution The patrol MPC slot held a 20-step GradientMPC started from the PID's rollout, which kept the follower within 29-32 m RMS of its slot on the protocol seeds. It now holds PatrolTwinOracle: the follower flies the lead's own heading autopilot on its own state, aimed at the slot, with small slot corrections and a turn feedforward, so in the slot it commands what the lead commands and the shared gust drops out of the relative position. A 30-step planner adds a residual to that law inside its rollout of step_env (40 Adam iterations, best iterate kept). Under the earlier floors the protocol cost falls from 3.02 to 0.0316 per step (-99%); the law alone gives 0.109. The slot floor, 18.6 m, was the old planner's own hold and 45x the new oracle's (0.41 m), so it moves to the 3 m relative-GPS resolution already in the params (slot_precision_floor). That changes the reward, and failure_cost with it (2.7e5 to 7.6e5 per step), so patrol and patrol_bearing_only, which share PatrolParams, become version 3. The 0.0087 rad heading floor is the AHRS resolution and stays. Re-measured under the new floor (scripts/measure_hold.py, seeds 0-1) the oracle holds 0.124 / 0.102 m of slot error and 1.63e-3 / 1.75e-3 rad of heading, so rho_floor = rho_floor_tracking = (0.1021 / 3)^2 + (1.631e-3 / 0.0087)^2 = 0.0363, from 1.03. The protocol gives 0.0371 per step against the PID's 265, NEA 1.000, zero trips, the audit's number to the last digit. Over ten seeds the MPC costs 14.2 per step against the PID's 299 and wins 10/10. patrol_bearing_only's MPC slot holds the same oracle reading the true state, labelled a full-state bound in its baselines_note: its NEA there measures information and control together. It gets a hold row like patrol's, so the protocol scores it after the capture (burn-in 100); its PID costs 199 there, and its MPC rows equal patrol's. The law carries the heading PID's unused lead_psi_prev field through unchanged, so the oracle's carry keeps its types from step to step and its jitted step compiles once per episode instead of twice (about 40 s each without a persistent compile cache). The heading PID never reads the field, so the actions are bit-identical to the twice-compiling version's (checked over 45 steps on patrol seed 0, and the re-recorded rows match it to the last digit). A fast test checks that the jit cache holds one entry. The protocol MPC now takes about 90 s per seed. The docs home page and the roadmap now say both patrol tasks ship an MPC, the reward-shaping intro lists them as -v3, and the proposal's fingerprint-gap list says patrol_bearing_only's new rows miss both the plane3d_heading gains and the patrol gains its PID reads. PHYSICS.md quotes the measured PID holds (44-68 m with full observation, 40-57 m bearing-only) instead of the old 229 m / 260 m claim, and the descriptions of baselines_note say it can also qualify a shipped baseline. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md # README.md # docs/baselines.md # docs/reward-shaping.md # src/target_gym/data/baseline_returns.json
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fifth per-plant upgrade from the oracle audit. The patrol oracle flies the lead's own autopilot, patrol_bearing_only gets an oracle, and both slot floors move to the GPS resolution (
-v3, approved).The patrol MPC slot held a 20-step GradientMPC started from the PID's
rollout, which kept the follower within 29-32 m RMS of its slot on the
protocol seeds. It now holds PatrolTwinOracle: the follower flies the
lead's own heading autopilot on its own state, aimed at the slot, with
small slot corrections and a turn feedforward, so in the slot it commands
what the lead commands and the shared gust drops out of the relative
position. A 30-step planner adds a residual to that law inside its rollout
of step_env (40 Adam iterations, best iterate kept). Under the earlier
floors the protocol cost falls from 3.02 to 0.0316 per step (-99%); the
law alone gives 0.109.
The slot floor, 18.6 m, was the old planner's own hold and 45x the new
oracle's (0.41 m), so it moves to the 3 m relative-GPS resolution already
in the params (slot_precision_floor). That changes the reward, and
failure_cost with it (2.7e5 to 7.6e5 per step), so patrol and
patrol_bearing_only, which share PatrolParams, become version 3. The
0.0087 rad heading floor is the AHRS resolution and stays. Re-measured
under the new floor (scripts/measure_hold.py, seeds 0-1) the oracle holds
0.124 / 0.102 m of slot error and 1.63e-3 / 1.75e-3 rad of heading, so
rho_floor = rho_floor_tracking = (0.1021 / 3)^2 + (1.631e-3 / 0.0087)^2
= 0.0363, from 1.03. The protocol gives 0.0371 per step against the PID's
265, NEA 1.000, zero trips, the audit's number to the last digit. Over ten
seeds the MPC costs 14.2 per step against the PID's 299 and wins 10/10.
patrol_bearing_only's MPC slot holds the same oracle reading the true
state, labelled a full-state bound in its baselines_note: its NEA there
measures information and control together. It gets a hold row like
patrol's, so the protocol scores it after the capture (burn-in 100); its
PID costs 199 there, and its MPC rows equal patrol's.
The law carries the heading PID's unused lead_psi_prev field through
unchanged, so the oracle's carry keeps its types from step to step and
its jitted step compiles once per episode instead of twice (about 40 s
each without a persistent compile cache). The heading PID never reads
the field, so the actions are bit-identical to the twice-compiling
version's (checked over 45 steps on patrol seed 0, and the re-recorded
rows match it to the last digit). A fast test checks that the jit cache
holds one entry. The protocol MPC now takes about 90 s per seed.
The docs home page and the roadmap now say both patrol tasks ship an
MPC, the reward-shaping intro lists them as -v3, and the proposal's
fingerprint-gap list says patrol_bearing_only's new rows miss both the
plane3d_heading gains and the patrol gains its PID reads. PHYSICS.md
quotes the measured PID holds (44-68 m with full observation, 40-57 m
bearing-only) instead of the old 229 m / 260 m claim, and the
descriptions of baselines_note say it can also qualify a shipped
baseline.
Checks
mkdocs build --strictpasses, and the contract, registry, version and docs tests pass (328 passed).GradientMPCenvironments, checked by importing the registry.Known gap (not fixed here)
The oracle reads the lead autopilot's gains (
pid_gains.jsonkeyplane3d_heading). Patrol's baseline fingerprint does not hash that key, and the plant itself already depends on it the same way. The fix belongs inprovenance.baseline_fingerprint, in its own PR.🤖 Generated with Claude Code