Conversation
Collaborator
|
Tell me when I should review. |
Author
|
I will test this on Monday with the flag run on jupiter |
Adds scripts/mass_convert_checkpoints.py and scripts/upload_to_hf.py for converting Megatron checkpoints from an already-running training job to sharded HuggingFace format and publishing them to a Hub repo, one branch per iteration, independent of the orchestrator's train/convert/eval chain. Both are idempotent and safe to re-run as new checkpoints appear. Supporting changes: - run_export.py / create_dummy_model.py: add --max-shard-size, threaded through the dummy reference model since the real export mirrors its shard layout. - validate_export.py: post-conversion sanity check, multilingual canonical prompts, writes validation.json alongside each checkpoint. - convert_and_validate_task.py: single-checkpoint Slurm task entry point used by mass_convert_checkpoints.py. - docs/bridge_eval_setup.md: document setup prerequisites and usage. Built and verified end to end publishing the OpenEuroLLM Prelude 9B checkpoints to openeurollm/prelude.
Also reverts an incorrect repo -> repository rename the doc-style autofixer made inside the literal --repo-id CLI flag name.
check-google-doc-style does a blind word-boundary substring rewrite (repo -> repository) that isn't code-aware, so it kept mangling the literal --repo-id CLI flag name into a nonexistent --repository-id. Wrapped both occurrences in google-doc-style-ignore/resume markers.
Adds --hf-repo-id (optional): skips converting a checkpoint that's already a complete branch on that HF repo, even if absent from this operator's own --output-dir. Lets multiple operators each convert into their own scratch dir (can't generally be shared across HPC accounts) against the same training run without redundantly reconverting what someone else already finished and uploaded. Also replaces the hardcoded Leonardo bind paths in the conversion job's singularity exec with a required --singularity-bind (repeatable, no default) so this isn't silently wrong on a different cluster.
Converting the 32B dense runs surfaced a set of problems the 9B never hit. Multiple runs in one Hub repository: - `--run-label` names outputs and branches `<label>_iter_NNNNNNN`. Without it two runs that reach the same iteration collide on one branch and discovery silently drops the second. - Manifests and job names are namespaced by label too; sharing an output dir previously meant the second run overwrote the first's `group_000.json`. - The uploader accepts a labelled dir name, so `<label>_iter_N` is discovered by the normal path rather than needing an explicit `--iters`. Robustness: - A failed `sbatch` no longer looks like a successful submission; the exit status is checked and the error reported. - Discovery tolerates an unreadable or dangling checkpoint symlink instead of aborting the whole pass. - An unreadable baked `tokenizer_model` path is a redirect reason like any other, rather than an uncaught PermissionError. Checkpoints can bake a path in another project's tree, readable only by named users. Conversion correctness: - The training layout is collapsed to a single rank before the model is built. Export runs one checkpoint per GPU, but the config came from the training args, so a custom `pipeline_model_parallel_layout` yielded a model holding only the first pipeline stage (5 of 64 layers) and the load then reported the other 59 shards missing. Parallelism-dependent comm optimisations are cleared with it. - Superseded MoE SM-count knobs are dropped; a newer core rejects a checkpoint that sets more than one of them. - `--vocab-size` overrides the reference tokenizer's length, needed when the tokenizer addresses more ids than the trained embedding has rows. - The image's own PYTHONPATH is preserved rather than replaced, and ordered so the image's matched megatron.core/megatron.bridge pair wins over a separate Bridge checkout. Publishing a tokenizer the model can represent: - Added tokens beyond `vocab_size` are dropped from the export and `pad_token` is repointed at `<unused_0>`, a real allocated row. `<eos>` would collide with the usual SFT recipe of masking labels where `input_ids == pad_token_id`, which would also mask every genuine end-of-sequence token. Only fires when the tokenizer overflows the embedding, so it is a no-op for the 9B, whose 262272 rows make its `<pad>` legal. - Validation asserts `len(tokenizer) <= embedding rows` and records the check in `validation.json`. Generation alone cannot catch this: the extra ids are special tokens the model never emits. Operational: - `--gpu-binding gres` for sites that reject `--gpus-per-task`, and each task then pins its own device, which that binding no longer does for it. - `--reservation` for test rounds on a development reservation. - Conversion runs as a child process. Bridge keeps live references to the loaded model, so its GPU memory survives `empty_cache()` and is still held when validation loads the export onto the same device. - `--prune-after-upload` deletes a local checkpoint once its branch is verified complete on the Hub; a full backfill is far larger than a project quota.
maxidl
force-pushed
the
exp_midahl00
branch
from
September 21, 2026 03:35
3823004 to
9510685
Compare
The exporter takes the architecture wholesale from the reference config, so converting against stock `Qwen/Qwen3-32B` produced a 151680-row embedding instead of 262144 while `config.json` said otherwise. Nothing in the conversion complained; only loading the export back caught it. The 9B works because its reference is an OpenEuroLLM config, not a stock Qwen one. `configs/openeurollm/oellm-32b` is the Qwen3-32B architecture with `vocab_size: 262144`, the runs' `padded_vocab_size`. Its reference tokenizer is the 256k tokenizer with `<pad>` removed: Bridge builds a tokenizer purely to compute the padded vocab when loading, so a 262145-entry tokenizer rounds the Megatron model up to 262272 and the exported tensors inherit that. It also carries no chat template, which keeps a Qwen chat protocol out of a base model's export. The documentation covers running this on JUPITER: checkpoints and training configs living on different file systems, `singularity` being Apptainer, the QOS limits, the development reservation, and generating the tokenizer's `tokenizer.json` once on a login node, since the Hub repository publishes only a SentencePiece `tokenizer.model` and `save_pretrained` rewrites the tracked metadata in passing.
maxidl
force-pushed
the
exp_midahl00
branch
from
September 21, 2026 03:48
9510685 to
938587f
Compare
`--prune-after-upload` only freed a checkpoint it uploaded in that pass. One completed by an earlier pass, by an interrupted round, or by another operator was skipped as already-complete and kept its disk forever. At ~64 GB per checkpoint that strands hundreds of GB over a large backfill, which is exactly the situation the flag exists to avoid. The branch is still re-verified against the Hub before anything is deleted.
Averaging a window of consecutive checkpoints gives a stronger, lower-variance quality readout than any single checkpoint, without paying for a decay run at each evaluation point. The merge arithmetic is NVIDIA's weighted_merge.py (Megatron-LM PR #5114); this adds the planner that decides which windows to merge and submits one job per merge. The stage sits before conversion because the tool reads and writes torch_dist and rejects anything else. That is also the correct side: TE FP8 _extra_state is copied from one chosen source rather than averaged and has no representation in safetensors, and a merged torch_dist checkpoint stays loadable for evaluation or resume. Each window writes a run-shaped output directory, so conversion is the existing call with the source run's training config and a distinct run label. Windows are laid out on the regular checkpoint grid. Segment-boundary saves such as iter_0084238 sit a few hundred iterations from their neighbour and would over-weight that point, so they are excluded as endpoints. The series is anchored at the newest checkpoint and walks back, and --merge-ignore-non-model- state is always passed because our checkpoints carry distributed optimizer state that the tool refuses to merge or silently copy.
The last checkpoint of a stable phase is usually off the regular grid, because training saves once more when the phase ends: prelude's is iter_0953312 against a 2400 grid. That checkpoint is exactly where a merge emulating the subsequent anneal has to end, so rejecting it as an endpoint ruled out the only comparison that matters. Endpoints now only have to exist, and the start iteration is derived by mirroring the tool's own backward walk, which preserves the target and keeps selected checkpoints at least one interval apart. An off-grid endpoint therefore selects exactly the requested number of inputs rather than over- or undershooting. Windows with too few checkpoints behind them are reported and skipped instead of silently producing a short merge. Grid membership still governs automatically generated endpoints, where an off-grid save would place two inputs a few hundred iterations apart.
A merge stands in for a particular anneal only if it ends where that anneal forked, spans the same tokens and follows the same decay shape. Restating those by hand is where this goes wrong, and the schedule is the quiet one: Nemotron merged with minus-sqrt because their decay was minus-sqrt, while prelude decays linear, so emulating it needs a uniform average instead. --match-anneal reads the anneal run's resolved config and derives all three: train_iters - lr_wsd_decay_iters gives the fork, the decay length gives the window, and lr_wsd_decay_style selects the merge style. Extra --window values still apply, so a sweep around the matched window is one invocation, and further anneal runs are one invocation each.
A series directory is named from the window and the style, so two sweeps that differ only in --min-iteration-interval resolve to the same path: a 483B span is 25 checkpoints at 2400, 13 at 4800 and 7 at 9600, and the last two would land on top of the first's output or be skipped as already present. Granularity is one of the axes worth sweeping -- WSM reports that merge duration matters more than checkpoint interval, which is a claim worth checking per model rather than assuming -- so --series-tag gives those sweeps distinct names.
Sweeps are routinely run side by side against one --output-dir -- different windows, schedules, checkpoint granularities or endpoints -- and the tracking file was rewritten as a whole JSON array on every submission. Two invocations interleaving their read-modify-write either lose records or abort the run mid-submission, which is how one sweep died after planning its windows. The file becomes append-only JSONL, one record per submitted job. A short O_APPEND write is atomic, so concurrent sweeps cannot corrupt it and no reader has to be consulted before writing. Verified with four concurrent writers: 800 records, none malformed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Standalone checkpoint conversion + HF Hub upload tooling
Summary
Adds a reusable, idempotent path for converting Megatron checkpoints from an
already-running training job to sharded HuggingFace format and publishing
them to a Hub repo (one branch per iteration), independent of the
orchestrator's train→convert→eval chain. Built and battle-tested end to end
converting and publishing all checkpoints from the OpenEuroLLM Prelude 9B
("baby") run (
openeurollm/prelude, 280+ checkpoints so far).What's included
--max-shard-sizeonrun_export.py/create_dummy_model.py: the realexport mirrors the dummy reference model's shard layout, so this has to be
threaded through both. Defaults to
5GB.validate_export.py: loads a converted checkpoint and runs acanonical-prompt, multilingual (36 EU languages) generation battery as a
post-conversion smoke test, writing
validation.jsonnext to thecheckpoint.
convert_and_validate_task.py: single-checkpoint Slurm task entry point(convert + validate in-process), driven by a JSON manifest + task index.
scripts/mass_convert_checkpoints.py: discovers unconverted checkpointsacross one or more source dirs, batches them into Slurm jobs respecting a
QOS's job cap, safe to re-run at any time (skips what's already done).
scripts/upload_to_hf.py: pushes each converted checkpoint to its own Hubbranch via
upload_large_folder(resumable); a branch only counts as doneonce it actually has the weight files, not just because the ref exists.
docs/bridge_eval_setup.md: new section documenting this path (setupprerequisites, example invocations, how it composes with a periodic
watcher for ongoing incremental use).
Why not part of the orchestrator chain
The chain in this doc assumes train/convert/eval are one managed pipeline
with a single output tree. This tooling targets a different, common case:
a training run is already in flight (or long finished), checkpoints exist
on disk, and the goal is "convert and publish everything, then keep up with
new ones as they land," independent of whether/how the run itself was
orchestrated.
Testing
pattern (conversion job →
sbatch --dependency=afterokupload job) forincremental use, and the
--dry-runcatch-up scan, all verified againstthe live
openeurollm/preluderepo.python3.11 -m py_compileon all changed/new files.Not included
new checkpoints appear (that's a separate PR to
oellm-monitoring, meantto run off-cluster so it isn't subject to login-node session limits).
uv.lockchanges (out of scope for this PR).