Skip to content

Add standalone checkpoint conversion + HF Hub upload tooling - #58

Draft
maxidl wants to merge 12 commits into
mainfrom
exp_midahl00
Draft

maxidl wants to merge 12 commits into
mainfrom
exp_midahl00

Conversation

@maxidl

@maxidl maxidl commented Jul 21, 2026

Copy link
Copy Markdown

Standalone checkpoint conversion + HF Hub upload tooling

Summary

Adds a reusable, idempotent path for converting Megatron checkpoints from an
already-running training job to sharded HuggingFace format and publishing
them to a Hub repo (one branch per iteration), independent of the
orchestrator's train→convert→eval chain. Built and battle-tested end to end
converting and publishing all checkpoints from the OpenEuroLLM Prelude 9B
("baby") run (openeurollm/prelude, 280+ checkpoints so far).

What's included

  • --max-shard-size on run_export.py / create_dummy_model.py: the real
    export mirrors the dummy reference model's shard layout, so this has to be
    threaded through both. Defaults to 5GB.
  • validate_export.py: loads a converted checkpoint and runs a
    canonical-prompt, multilingual (36 EU languages) generation battery as a
    post-conversion smoke test, writing validation.json next to the
    checkpoint.
  • convert_and_validate_task.py: single-checkpoint Slurm task entry point
    (convert + validate in-process), driven by a JSON manifest + task index.
  • scripts/mass_convert_checkpoints.py: discovers unconverted checkpoints
    across one or more source dirs, batches them into Slurm jobs respecting a
    QOS's job cap, safe to re-run at any time (skips what's already done).
  • scripts/upload_to_hf.py: pushes each converted checkpoint to its own Hub
    branch via upload_large_folder (resumable); a branch only counts as done
    once it actually has the weight files, not just because the ref exists.
  • docs/bridge_eval_setup.md: new section documenting this path (setup
    prerequisites, example invocations, how it composes with a periodic
    watcher for ongoing incremental use).

Why not part of the orchestrator chain

The chain in this doc assumes train/convert/eval are one managed pipeline
with a single output tree. This tooling targets a different, common case:
a training run is already in flight (or long finished), checkpoints exist
on disk, and the goal is "convert and publish everything, then keep up with
new ones as they land," independent of whether/how the run itself was
orchestrated.

Testing

  • End-to-end on Leonardo: full backfill of 280+ checkpoints, dependency
    pattern (conversion job → sbatch --dependency=afterok upload job) for
    incremental use, and the --dry-run catch-up scan, all verified against
    the live openeurollm/prelude repo.
  • python3.11 -m py_compile on all changed/new files.

Not included

  • A periodic watcher/scheduler that invokes this tooling automatically as
    new checkpoints appear (that's a separate PR to oellm-monitoring, meant
    to run off-cluster so it isn't subject to login-node session limits).
  • Any project-wide uv.lock changes (out of scope for this PR).

@kpoeppel

Copy link
Copy Markdown
Collaborator

Tell me when I should review.

@maxidl

maxidl commented Aug 28, 2026

Copy link
Copy Markdown
Author

I will test this on Monday with the flag run on jupiter

maxidl and others added 5 commits September 20, 2026 19:05
Adds scripts/mass_convert_checkpoints.py and scripts/upload_to_hf.py for
converting Megatron checkpoints from an already-running training job to
sharded HuggingFace format and publishing them to a Hub repo, one branch
per iteration, independent of the orchestrator's train/convert/eval chain.
Both are idempotent and safe to re-run as new checkpoints appear.

Supporting changes:
- run_export.py / create_dummy_model.py: add --max-shard-size, threaded
  through the dummy reference model since the real export mirrors its
  shard layout.
- validate_export.py: post-conversion sanity check, multilingual canonical
  prompts, writes validation.json alongside each checkpoint.
- convert_and_validate_task.py: single-checkpoint Slurm task entry point
  used by mass_convert_checkpoints.py.
- docs/bridge_eval_setup.md: document setup prerequisites and usage.

Built and verified end to end publishing the OpenEuroLLM Prelude 9B
checkpoints to openeurollm/prelude.
Also reverts an incorrect repo -> repository rename the doc-style
autofixer made inside the literal --repo-id CLI flag name.
check-google-doc-style does a blind word-boundary substring rewrite
(repo -> repository) that isn't code-aware, so it kept mangling the
literal --repo-id CLI flag name into a nonexistent --repository-id.
Wrapped both occurrences in google-doc-style-ignore/resume markers.
Adds --hf-repo-id (optional): skips converting a checkpoint that's
already a complete branch on that HF repo, even if absent from this
operator's own --output-dir. Lets multiple operators each convert into
their own scratch dir (can't generally be shared across HPC accounts)
against the same training run without redundantly reconverting what
someone else already finished and uploaded.

Also replaces the hardcoded Leonardo bind paths in the conversion job's
singularity exec with a required --singularity-bind (repeatable, no
default) so this isn't silently wrong on a different cluster.
Converting the 32B dense runs surfaced a set of problems the 9B never hit.

Multiple runs in one Hub repository:
- `--run-label` names outputs and branches `<label>_iter_NNNNNNN`. Without it
  two runs that reach the same iteration collide on one branch and discovery
  silently drops the second.
- Manifests and job names are namespaced by label too; sharing an output dir
  previously meant the second run overwrote the first's `group_000.json`.
- The uploader accepts a labelled dir name, so `<label>_iter_N` is discovered
  by the normal path rather than needing an explicit `--iters`.

Robustness:
- A failed `sbatch` no longer looks like a successful submission; the exit
  status is checked and the error reported.
- Discovery tolerates an unreadable or dangling checkpoint symlink instead of
  aborting the whole pass.
- An unreadable baked `tokenizer_model` path is a redirect reason like any
  other, rather than an uncaught PermissionError. Checkpoints can bake a path
  in another project's tree, readable only by named users.

Conversion correctness:
- The training layout is collapsed to a single rank before the model is built.
  Export runs one checkpoint per GPU, but the config came from the training
  args, so a custom `pipeline_model_parallel_layout` yielded a model holding
  only the first pipeline stage (5 of 64 layers) and the load then reported
  the other 59 shards missing. Parallelism-dependent comm optimisations are
  cleared with it.
- Superseded MoE SM-count knobs are dropped; a newer core rejects a checkpoint
  that sets more than one of them.
- `--vocab-size` overrides the reference tokenizer's length, needed when the
  tokenizer addresses more ids than the trained embedding has rows.
- The image's own PYTHONPATH is preserved rather than replaced, and ordered so
  the image's matched megatron.core/megatron.bridge pair wins over a separate
  Bridge checkout.

Publishing a tokenizer the model can represent:
- Added tokens beyond `vocab_size` are dropped from the export and `pad_token`
  is repointed at `<unused_0>`, a real allocated row. `<eos>` would collide
  with the usual SFT recipe of masking labels where `input_ids ==
  pad_token_id`, which would also mask every genuine end-of-sequence token.
  Only fires when the tokenizer overflows the embedding, so it is a no-op for
  the 9B, whose 262272 rows make its `<pad>` legal.
- Validation asserts `len(tokenizer) <= embedding rows` and records the check
  in `validation.json`. Generation alone cannot catch this: the extra ids are
  special tokens the model never emits.

Operational:
- `--gpu-binding gres` for sites that reject `--gpus-per-task`, and each task
  then pins its own device, which that binding no longer does for it.
- `--reservation` for test rounds on a development reservation.
- Conversion runs as a child process. Bridge keeps live references to the
  loaded model, so its GPU memory survives `empty_cache()` and is still held
  when validation loads the export onto the same device.
- `--prune-after-upload` deletes a local checkpoint once its branch is
  verified complete on the Hub; a full backfill is far larger than a project
  quota.
The exporter takes the architecture wholesale from the reference config, so
converting against stock `Qwen/Qwen3-32B` produced a 151680-row embedding
instead of 262144 while `config.json` said otherwise. Nothing in the
conversion complained; only loading the export back caught it. The 9B works
because its reference is an OpenEuroLLM config, not a stock Qwen one.

`configs/openeurollm/oellm-32b` is the Qwen3-32B architecture with
`vocab_size: 262144`, the runs' `padded_vocab_size`. Its reference tokenizer
is the 256k tokenizer with `<pad>` removed: Bridge builds a tokenizer purely
to compute the padded vocab when loading, so a 262145-entry tokenizer rounds
the Megatron model up to 262272 and the exported tensors inherit that. It also
carries no chat template, which keeps a Qwen chat protocol out of a base
model's export.

The documentation covers running this on JUPITER: checkpoints and training
configs living on different file systems, `singularity` being Apptainer, the
QOS limits, the development reservation, and generating the tokenizer's
`tokenizer.json` once on a login node, since the Hub repository publishes only
a SentencePiece `tokenizer.model` and `save_pretrained` rewrites the tracked
metadata in passing.
`--prune-after-upload` only freed a checkpoint it uploaded in that pass. One
completed by an earlier pass, by an interrupted round, or by another operator
was skipped as already-complete and kept its disk forever. At ~64 GB per
checkpoint that strands hundreds of GB over a large backfill, which is exactly
the situation the flag exists to avoid. The branch is still re-verified against
the Hub before anything is deleted.
Averaging a window of consecutive checkpoints gives a stronger, lower-variance
quality readout than any single checkpoint, without paying for a decay run at
each evaluation point. The merge arithmetic is NVIDIA's weighted_merge.py
(Megatron-LM PR #5114); this adds the planner that decides which windows to
merge and submits one job per merge.

The stage sits before conversion because the tool reads and writes torch_dist
and rejects anything else. That is also the correct side: TE FP8 _extra_state is
copied from one chosen source rather than averaged and has no representation in
safetensors, and a merged torch_dist checkpoint stays loadable for evaluation or
resume. Each window writes a run-shaped output directory, so conversion is the
existing call with the source run's training config and a distinct run label.

Windows are laid out on the regular checkpoint grid. Segment-boundary saves such
as iter_0084238 sit a few hundred iterations from their neighbour and would
over-weight that point, so they are excluded as endpoints. The series is
anchored at the newest checkpoint and walks back, and --merge-ignore-non-model-
state is always passed because our checkpoints carry distributed optimizer
state that the tool refuses to merge or silently copy.
The last checkpoint of a stable phase is usually off the regular grid, because
training saves once more when the phase ends: prelude's is iter_0953312 against
a 2400 grid. That checkpoint is exactly where a merge emulating the subsequent
anneal has to end, so rejecting it as an endpoint ruled out the only comparison
that matters.

Endpoints now only have to exist, and the start iteration is derived by
mirroring the tool's own backward walk, which preserves the target and keeps
selected checkpoints at least one interval apart. An off-grid endpoint therefore
selects exactly the requested number of inputs rather than over- or
undershooting. Windows with too few checkpoints behind them are reported and
skipped instead of silently producing a short merge.

Grid membership still governs automatically generated endpoints, where an
off-grid save would place two inputs a few hundred iterations apart.
A merge stands in for a particular anneal only if it ends where that anneal
forked, spans the same tokens and follows the same decay shape. Restating those
by hand is where this goes wrong, and the schedule is the quiet one: Nemotron
merged with minus-sqrt because their decay was minus-sqrt, while prelude decays
linear, so emulating it needs a uniform average instead.

--match-anneal reads the anneal run's resolved config and derives all three:
train_iters - lr_wsd_decay_iters gives the fork, the decay length gives the
window, and lr_wsd_decay_style selects the merge style. Extra --window values
still apply, so a sweep around the matched window is one invocation, and further
anneal runs are one invocation each.
A series directory is named from the window and the style, so two sweeps that
differ only in --min-iteration-interval resolve to the same path: a 483B span is
25 checkpoints at 2400, 13 at 4800 and 7 at 9600, and the last two would land on
top of the first's output or be skipped as already present.

Granularity is one of the axes worth sweeping -- WSM reports that merge duration
matters more than checkpoint interval, which is a claim worth checking per model
rather than assuming -- so --series-tag gives those sweeps distinct names.
Sweeps are routinely run side by side against one --output-dir -- different
windows, schedules, checkpoint granularities or endpoints -- and the tracking
file was rewritten as a whole JSON array on every submission. Two invocations
interleaving their read-modify-write either lose records or abort the run
mid-submission, which is how one sweep died after planning its windows.

The file becomes append-only JSONL, one record per submitted job. A short
O_APPEND write is atomic, so concurrent sweeps cannot corrupt it and no reader
has to be consulted before writing. Verified with four concurrent writers: 800
records, none malformed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants