Repository navigation
Speed up sweep composition - #63
Merged
Merged
Conversation
Mirrors the change in hydra_staged_sweep into the vendored copy. Building a sweep composes the same config tree once per sweep point, and Hydra keeps nothing between compose() calls: every point re-read and re-parsed the same YAML, re-walked the same defaults list, re-merged the same configs and re-parsed the same interpolation and override strings. hydra_staged_sweep/config/cache.py adds six in-memory caches over Hydra and OmegaConf -- parsed config files, config-group lookups, merged defaults lists, interpolation parse trees, override parses, and a libyaml-backed YAML loader. Entries carry a stat() fingerprint and are revalidated on every lookup, so editing a config invalidates exactly what depends on it. They install when hydra_staged_sweep/config/loader.py is imported; HYDRA_STAGED_SWEEP_CACHE=0 turns them off. Staged sweeps also stop recomposing siblings: resolve_sweep_with_dag reuses the sibling's already-resolved JobPlan config, builds the sibling context once per stage chain, and merges it into the composed config in one step instead of flattening it into several hundred ++key=value overrides for Hydra to parse and apply one at a time. job_parameters still records those overrides, so a job stays reproducible from the command line. run_autoexp.py --dry-run on the 300-job titan_qwen3_depth_width_scaling sweep goes from 469s to 82s (5.75x), with all 300 rendered sbatch scripts byte-identical to before.
Mirrors the change in hydra_staged_sweep into the vendored copy, and fans out script rendering in the orchestrator the same way. Sweep resolution is pure CPU: composing, resolving interpolations and building dataclasses. A sweep splits into independent dependency chains -- a stable stage and the cooldowns branching off it must run in order, but separate chains share nothing -- so resolve_sweep_with_dag hands the chains to a fork pool. Forking means workers inherit the loaded modules, the registered resolvers and the warm config caches, so only each chain and its results cross the process boundary. compoconf configs cannot pickle themselves -- ConfigInterface.__reduce__ reduces to (cls, (), state), so unpickling calls cls() with no arguments, which raises for any config with required fields, and its asdict state flattens nested configs into plain dicts. Plans are pickled by __dict__ instead. Two more caches, both of which the pool then multiplies: the Defaults List is memoized on the overrides that can actually select a config group, and CachingConfigRepository no longer deep-copies the whole repository on every composition. Filter evaluation builds its context lazily, since a filter that is already a bool never reads it. HYDRA_STAGED_SWEEP_WORKERS pins the pool size; 1 keeps everything in-process. run_autoexp.py --dry-run on the 300-job titan_qwen3_depth_width_scaling sweep goes from 82s to 21s -- 470s before any of this work -- with all 300 rendered sbatch scripts still byte-identical.
compoconf 0.2.1 fixes ConfigInterface pickling: configs now use the default dataclass reduction, which restores state onto a new instance without calling __init__ and keeps nested configs typed. That is exactly what _PlanPickler did by hand, so plans can cross the process boundary on the pool's own pickler and _restore_config / _PlanPickler / _dumps all go away. The floor is compoconf>=0.2.2, which is the first release carrying both that fix and the tests covering it. 0.2.1 is the oldest release that works; 0.2.0 and earlier cannot unpickle a worker's results at all, because cls() with no arguments raises for configs with required fields, and nested configs come back as plain dicts. While here, the filter context becomes a small class instead of a closure over loop variables, and the dead context/skip_point assignments in the no-sweep branch go -- they were unreachable leftovers, and moving the resolve loop into _resolve_chain made that visible.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.