Skip to content

Speed up sweep composition - #63

Merged
kpoeppel merged 4 commits into
mainfrom
speed-up-sweep-composition
Aug 2, 2026
Merged

kpoeppel merged 4 commits into
mainfrom
speed-up-sweep-composition

Conversation

@kpoeppel

@kpoeppel kpoeppel commented Aug 2, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

Mirrors the change in hydra_staged_sweep into the vendored copy.

Building a sweep composes the same config tree once per sweep point, and Hydra
keeps nothing between compose() calls: every point re-read and re-parsed the
same YAML, re-walked the same defaults list, re-merged the same configs and
re-parsed the same interpolation and override strings.

hydra_staged_sweep/config/cache.py adds six in-memory caches over Hydra and
OmegaConf -- parsed config files, config-group lookups, merged defaults lists,
interpolation parse trees, override parses, and a libyaml-backed YAML loader.
Entries carry a stat() fingerprint and are revalidated on every lookup, so
editing a config invalidates exactly what depends on it. They install when
hydra_staged_sweep/config/loader.py is imported; HYDRA_STAGED_SWEEP_CACHE=0
turns them off.

Staged sweeps also stop recomposing siblings: resolve_sweep_with_dag reuses
the sibling's already-resolved JobPlan config, builds the sibling context once
per stage chain, and merges it into the composed config in one step instead of
flattening it into several hundred ++key=value overrides for Hydra to parse
and apply one at a time. job_parameters still records those overrides, so a
job stays reproducible from the command line.

run_autoexp.py --dry-run on the 300-job titan_qwen3_depth_width_scaling
sweep goes from 469s to 82s (5.75x), with all 300 rendered sbatch scripts
byte-identical to before.
Mirrors the change in hydra_staged_sweep into the vendored copy, and fans out
script rendering in the orchestrator the same way.

Sweep resolution is pure CPU: composing, resolving interpolations and building
dataclasses. A sweep splits into independent dependency chains -- a stable
stage and the cooldowns branching off it must run in order, but separate chains
share nothing -- so resolve_sweep_with_dag hands the chains to a fork pool.
Forking means workers inherit the loaded modules, the registered resolvers and
the warm config caches, so only each chain and its results cross the process
boundary.

compoconf configs cannot pickle themselves -- ConfigInterface.__reduce__
reduces to (cls, (), state), so unpickling calls cls() with no arguments, which
raises for any config with required fields, and its asdict state flattens
nested configs into plain dicts. Plans are pickled by __dict__ instead.

Two more caches, both of which the pool then multiplies: the Defaults List is
memoized on the overrides that can actually select a config group, and
CachingConfigRepository no longer deep-copies the whole repository on every
composition. Filter evaluation builds its context lazily, since a filter that
is already a bool never reads it.

HYDRA_STAGED_SWEEP_WORKERS pins the pool size; 1 keeps everything in-process.

run_autoexp.py --dry-run on the 300-job titan_qwen3_depth_width_scaling sweep
goes from 82s to 21s -- 470s before any of this work -- with all 300 rendered
sbatch scripts still byte-identical.
compoconf 0.2.1 fixes ConfigInterface pickling: configs now use the default
dataclass reduction, which restores state onto a new instance without calling
__init__ and keeps nested configs typed. That is exactly what _PlanPickler did
by hand, so plans can cross the process boundary on the pool's own pickler and
_restore_config / _PlanPickler / _dumps all go away.

The floor is compoconf>=0.2.2, which is the first release carrying both that
fix and the tests covering it. 0.2.1 is the oldest release that works; 0.2.0
and earlier cannot unpickle a worker's results at all, because cls() with no
arguments raises for configs with required fields, and nested configs come back
as plain dicts.

While here, the filter context becomes a small class instead of a closure over
loop variables, and the dead context/skip_point assignments in the no-sweep
branch go -- they were unreachable leftovers, and moving the resolve loop into
_resolve_chain made that visible.
@kpoeppel
kpoeppel merged commit de79aaa into main Aug 2, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant