@@ -108,7 +108,7 @@ Each entry calls `pipeline.load_lora_weights(<lora_id>, adapter_name=<name>)`. A
108108- ` --cpu-offload {model, group} ` — ` model ` uses ` enable_model_cpu_offload ` , ` group ` uses
109109 ` enable_group_offload(offload_type="leaf_level", use_stream=True) ` . Use ` group ` to fit a 9B+ model on a single
110110 A100. Onload target device comes from ` --device-map ` (must be a plain device string in this case).
111- - ` --attention-backend {default, flash_hub, flash_varlen_hub, flash_4_hub, sage_hub} ` — hub-hosted kernels,
111+ - ` --attention-backend {default, flash_hub, flash_varlen_hub, flash_4_hub, sage_hub, sage_blackwell_hub } ` — hub-hosted kernels,
112112 auto-downloaded on first use. Failures (kernel not available, CUDA arch mismatch, network) raise a clear
113113 ` SystemExit ` listing the alternatives instead of silently reverting to the default. Only supported on
114114 transformer-based pipelines; UNet pipelines get a ` logger.warning ` and the flag is ignored.
@@ -151,8 +151,8 @@ user also explicitly asked for a local target via `--output`.
151151## Remote execution (` --remote ` )
152152
153153Add ` --remote ` to run the same call inside a [ Hugging Face Sandbox] ( https://huggingface.co/docs/huggingface_hub/en/guides/sandbox )
154- — an isolated cloud VM (built on HF Jobs) the CLI drives over HTTP: it uploads inputs, installs deps, runs
155- the pipeline, downloads outputs, then terminates the sandbox.
154+ — an isolated cloud VM (built on HF Jobs) the CLI drives over HTTP: it uploads inputs, runs the pipeline,
155+ downloads outputs, then terminates the sandbox.
156156
157157``` bash
158158diffusers-cli run \
@@ -167,12 +167,14 @@ What happens:
167167
1681681 . Your HF token is picked up (from ` --token ` or your login) and forwarded into the sandbox as ` HF_TOKEN ` .
1691692 . ` --pipeline-kwargs ` are parsed locally so JSON errors fail fast (no wasted sandbox time).
170- 3 . A dedicated sandbox is created on ` --flavor ` from a pytorch image
171- (` pytorch/pytorch:2.10.0-cuda12.8-cudnn9-runtime ` by default) that already has torch + CUDA. Any local
172- file paths in ` --pipeline-kwargs ` are uploaded into the sandbox under ` /tmp/diffusers-cli/inputs/<run_id>/ `
173- via native file transfer (no bucket), and the JSON paths are rewritten to point at them.
174- 4 . The small Python deps (` diffusers ` , ` accelerate ` , ` transformers ` , ` safetensors ` , ` sentencepiece ` , ` ftfy ` )
175- are installed with ` uv pip install --system ` . Output (install + run) streams live to your terminal.
170+ 3 . A dedicated sandbox is created on ` --flavor ` from ` diffusers/diffusers-cli-cuda:latest ` by default — a
171+ prebuilt image (` docker/diffusers-cli-cuda/ ` in this repo, rebuilt nightly) that already ships torch,
172+ CUDA, ` diffusers ` , and the rest of the CLI's deps. Any local file paths in ` --pipeline-kwargs ` are
173+ uploaded into the sandbox under ` /tmp/diffusers-cli/inputs/<run_id>/ ` via native file transfer (no
174+ bucket), and the JSON paths are rewritten to point at them.
175+ 4 . No dependency install runs on the default image — that step is skipped unless you passed ` --dependencies `
176+ (only the extras are installed then) or pointed ` --image ` somewhere else (the full set is installed).
177+ Output streams live to your terminal.
1761785 . The sandbox CLI writes outputs to ` /tmp/diffusers-cli/outputs/ ` ; the CLI downloads every artifact back into
177179 the local target (see [ ` --push-to ` ] ( #-push-to ) for when the download is skipped).
1781806 . The sandbox is terminated (unless ` --keep-alive ` /` --sandbox-id ` ), and the wallclock ` run_seconds ` for the
@@ -182,12 +184,13 @@ Flags:
182184
183185- ` --flavor <name> ` — sandbox hardware (e.g. ` a10g-small ` , ` a100-large ` , ` 4xa100-large ` ).
184186- ` --timeout <duration> ` — max wallclock for the run command inside the sandbox (e.g. ` 30m ` , ` 2h ` ). Defaults to ` 10m ` .
185- - ` --dependencies <pkg> ` — extra pip deps (repeatable). Appends to the defaults .
187+ - ` --dependencies <pkg> ` — extra pip deps (repeatable), installed on top of whatever the image ships .
186188- ` --namespace <name> ` — create the sandbox under a different account.
187189- ` --push-to <bucket> ` — see [ ` --push-to ` ] ( #-push-to ) above. The upload runs inside the sandbox; an explicit
188190 value with no ` --output ` makes the bucket the sole destination and skips the local download.
189- - ` --image <ref> ` — override the sandbox image. Must ship torch + CUDA; the CLI installs the small Python
190- deps on top via ` uv pip install --system ` . Useful for pinning a specific torch or bundling extra system libs.
191+ - ` --image <ref> ` — override the sandbox image. Must ship torch + CUDA; the CLI then installs the small
192+ Python deps on top via ` uv pip install --system ` on every cold sandbox, which the default image avoids.
193+ Useful for pinning a specific torch or bundling extra system libs.
191194- ` --volume <bucket-id>[:<mount-path>] ` — mount an HF storage bucket into the sandbox as a read-write directory.
192195 Repeatable. Default mount is ` /mnt/buckets/<bucket-id> ` . Reference mounted files from ` --pipeline-kwargs `
193196 like any other local path — no upload happens, the container reads straight from the FUSE mount. Applied
0 commit comments