pubflow is a lightweight workflow manager for publishing ESGF datasets from mapfiles. It provides:
-
A persistent DuckDB database for tracking campaigns, datasets, files, publication attempts, and archival state.
-
A Typer-based command-line interface for managing publication workflows.
-
Configurable batch publication, retries, and dry-run support.
-
Multiple ESGF publisher configuration profiles for different deployment environments.
-
CSV and Grist status export capabilities.
-
A decoupled archival workflow for computing centres where the publisher lacks direct write access to the final storage location.
The workflow is designed to keep orchestration concerns separate from the underlying esgpublish application and from
the computing centres hosting the published data.
The workflow is divided into registration, publication, status export, and archival stages:
Publisher VM
|
+----------------+----------------+
| |
v v
Register Publish
| |
v v
DuckDB <---------------- Publication status
|
+-----+------+
| |
v v
Export Grist
|
v
Dashboard
Archival is deliberately decoupled:
Publisher VM
|
v
Generate archive tasks
|
v
archive_tasks.csv
|
| portable transfer
v
Computing Centre
|
v
bin/archive.py
|
v
archive_results.csv
Note: The publisher does not need write access to the final archive filesystem.
The project is a Python application that exposes the pubflow command.
After installing the package in the appropriate environment:
pubflow --help
| Command | Description |
|---|---|
| pubflow dataset register | Register datasets and files from mapfiles |
| pubflow publication run | Publish datasets |
| pubflow publication diagnose | Re-run failed datasets individually and capture server diagnostics |
| pubflow stac clean | Preview or remove publisher-generated STAC JSON dumps |
| pubflow dataset validate | Validate registered datasets |
| pubflow report export | Export database state to CSV |
| pubflow grist | Synchronize workflow status with Grist |
| pubflow archive generate | Generate portable archive tasks |
The 0.2 command line follows pubflow <subject> <action>. The former
top-level action commands remain available as hidden, deprecated aliases for
the 0.2 release, except pubflow archive: the archive name is now the
command group, so task generation uses pubflow archive generate.
Note: Running
pubflowwithout a command displays command help.
Pubflow uses several configuration files, each with a deliberately separate responsibility.
config/
├── campaigns.yml
└── publisher.yml
~/.esg/
├── esg.yaml.EASTINT
├── esg.yaml.WESTINT
├── esg.yaml.EAST
└── esg.yaml.WEST
Campaigns are defined in:
config/campaigns.yml
A campaign contains the project/activity metadata and the locations required by the workflow:
campaigns:
tipmip-cnrm:
project: CMIP6Plus
activity: TIPMIP
institution: CNRM-CERFACS
mapfile_root: /modfs/esgf/topublish/CNRM-CERFACS/.mapfiles
archive:
enabled: true
root: /mnt/scality/WCRP/CMIP6Plus/TIPMIP/CNRM-CERFACS
depth: experiment_id
Campaigns therefore retain both:
-
project -
activity
These fields are also exported to Grist and are used by the dashboard to filter campaigns.
Publisher execution is configured separately in:
config/publisher.yml
For example:
publisher:
executable: esgpublish
arguments:
- --no-xarray
batch:
size: 50
execution:
dry_run: false
logging:
directory: logs
retry:
enabled: true
max_attempts: 3
mapfile_path_mappings:
- from: /ccc/work/cont003/cmip6/cmip6
to: /mnt/tgcc/
esg:
config:
active: EAST-int
profiles:
EAST-int:
path: ~/.esg/esg.yaml.EASTINT
WEST-int:
path: ~/.esg/esg.yaml.WESTINT
EAST-prod:
path: ~/.esg/esg.yaml.EAST
WEST-prod:
path: ~/.esg/esg.yaml.WEST
The publisher configuration intentionally contains only the generic esgpublish execution settings.
The actual ESG publisher configuration is selected through an ESG configuration profile.
When a dataset is published, Pubflow effectively constructs:
esgpublish \
--no-xarray \
--config ~/.esg/esg.yaml.EASTINT \
--map <mapfile>
This keeps the publisher command configuration independent from the campaign definitions.
The profile system allows the same Pubflow installation to target different ESG publisher environments.
For example:
EAST-int
WEST-int
EAST-prod
WEST-prod
The active profile is selected in publisher.yml:
esg:
config:
active: EAST-int
Before publication, Pubflow verifies that:
-
The selected profile exists.
-
The configured ESG configuration file exists.
An invalid profile or missing configuration therefore fails early rather than producing a less useful publisher error.
Note: Pubflow does not currently interpret or modify the contents of the
esg.yamlfiles. They remain configuration files owned by the ESG publisher environment.
Authentication is deliberately handled by esgpublish.
The ESG publisher stores its authentication token in the location configured by the selected ESG configuration, for example:
~/.esgf-publisher.json
When the token expires, esgpublish can initiate the EGI Check-in authentication flow and display the URL that must be
visited to renew authentication.
Pubflow does not currently attempt to manage or renew these tokens.
This is intentional:
-
esgpublishremains responsible for authentication. -
Pubflow does not need to understand the token format.
-
Authentication behaviour remains consistent with standalone
esgpublish. -
Manual renewal remains possible when required.
Operational note: A publication process may pause while waiting for manual authentication. For long-running campaigns, running Pubflow inside
tmuxor another persistent terminal session is recommended.
Pubflow uses DuckDB to maintain persistent workflow state.
The database tracks:
-
Campaigns
-
Datasets
-
Files
-
Publication attempts
-
Archive status
The schema is defined in:
db/schema.sql
Initialize the database with:
python bin/init_db.py
Load campaign definitions with:
python workflow/campaign.py
Registration scans a campaign's mapfile directory and registers datasets and their files in DuckDB.
pubflow dataset register tipmip-cnrm
Example output:
Found 2446 mapfiles
Registered CMIP6Plus.... (12 files)
Registered CMIP6Plus.... (8 files)
...
Completed: 2445 succeeded, 1 failed
Note: Registration does not publish anything.
Registration is therefore safe to run before a publication campaign begins.
Publish datasets for a campaign:
pubflow publication run tipmip-cnrm
Use a limit when testing:
pubflow publication run tipmip-cnrm --limit 10
Publication is performed in batches, with one esgpublish directory invocation
per batch:
pubflow publication run tipmip-cnrm --batch-size 50
The default batch size is 50. Set --batch-size 1 to retain the previous
one-invocation-per-dataset behavior.
PUB_STATUS=PASS and PUB_STATUS=FAIL messages are recorded per dataset. If
the publisher stops at the first failure, mapfiles it did not reach remain
PENDING and are selected for the next batch. The executor tracks each
reported publication attempt and updates the corresponding dataset state in
DuckDB.
If an entire batch returns no recognizable PUB_STATUS, Pubflow retries its
first dataset alone. If the isolated retry is also status-less, the attempt is
recorded as NO_STATUS, the dataset remains PENDING, and it is deferred for
the remainder of the current run. Publication then continues with the other
pending datasets. The run summary and log list every deferred dataset, along
with the publisher exit code and output length.
Retries can be configured in publisher.yml:
publisher:
retry:
enabled: true
max_attempts: 3
The publisher execution mode can be controlled through:
publisher:
execution:
dry_run: false
Each publication run receives a unique run identifier and an associated log file.
Run information includes:
-
Campaign
-
Run ID
-
Start time
-
End time
-
Batch information
-
Mapfile path mappings
-
Dataset successes
-
Dataset failures
-
Exit codes
-
Error messages
-
Publication summary
The configured log directory is:
publisher:
logging:
directory: logs
Publication failures are also persisted in DuckDB and exported to Grist.
Run diagnostics after a publication campaign to retry its currently failed datasets one at a time and capture the detailed response returned by the transaction service:
pubflow publication diagnose tipmip-cnrm
This is a real publication attempt, not a dry run. A dataset which now emits
PUB_STATUS=PASS is recorded as SUCCESS; a dataset which still fails remains
FAILED. If the publisher emits no recognizable status, Pubflow records an
execution error without changing the dataset's publication status.
Each run creates a directory below
logs/<campaign>/diagnostics/<diagnostic-run-id>/ containing one full publisher
log per dataset and a concise diagnostics.csv. Results are also appended to
the DuckDB diagnostic_attempts table. Existing publication records are not
removed or replaced.
Pubflow publication runs force temporary STAC generation with --save-stac.
Diagnostics additionally enable --verbose and validate the generated item
locally for basic STAC structure, WGS84
geometry, and every schema declared in stac_extensions. Schemas are fetched
once and cached in memory for the diagnostic run. Field-level local validation
errors take precedence over a generic EAST response and all errors are written
to the dataset log.
The publisher writes items into an isolated staging directory. Pubflow retains
only items with an explicit PUB_STATUS=FAIL in stac-items/; successful and
unreached items are discarded. If a previously failed dataset later succeeds,
its stale retained item is removed. The destination can be changed in
config/publisher.yml:
publisher:
stac_items:
directory: stac-items
It can also be overridden with PUBFLOW_STAC_ITEMS_DIR. To retain an additional
per-run copy alongside the diagnostic logs, including a recovered item:
pubflow publication diagnose tipmip-cnrm --persist-stac-item
Older or manually generated STAC dumps can be cleaned without shell wildcard expansion. The command previews matches by default and verifies the JSON content and item ID before treating a file as a generated STAC Item:
pubflow stac clean \
--directory /home/esguser/esgf-publisher-workflow \
--pattern 'CMIP6.ScenarioMIP.IPSL.IPSL-CM*.json'
After reviewing the preview, repeat with --delete:
pubflow stac clean \
--directory /home/esguser/esgf-publisher-workflow \
--pattern 'CMIP6.ScenarioMIP.IPSL.IPSL-CM*.json' \
--delete
Cleanup is limited to the selected directory and does not recurse into subdirectories.
The installed EAST publisher must support saving STAC from its EGI transaction client. Detailed transaction service output remains useful for authorization, conflict, and other server-only failures, but it is no longer required for field-level schema diagnostics.
Grist synchronization is enabled by default. Use --no-sync-grist to keep the
run local, or create the Diagnostics table described below before enabling
synchronization. A Grist failure is reported as a warning and does not discard
the DuckDB, CSV, or log results.
Useful limiting form for the first production run:
pubflow publication diagnose tipmip-cnrm --limit 5 --no-sync-grist
Validate registered datasets without triggering publication:
pubflow dataset validate tipmip-cnrm
Validation does not invoke the publication workflow.
Export the current database state to CSV:
pubflow report export --output campaign_status.csv
CSV export is useful for inspecting workflow state outside the Pubflow environment or for downstream processing.
Pubflow can synchronize campaign, dataset, and failure information with Grist:
pubflow grist sync
The Grist document contains three main workflow tables and an optional diagnostics table:
Campaigns
Datasets
Failures
Diagnostics
The campaign table contains:
| Field | Description |
|---|---|
| campaign | Campaign name |
| project | Project identifier |
| activity | Activity identifier |
| total | Total datasets |
| published | Successfully published datasets |
| failed | Failed datasets |
| pending | Datasets still pending |
The dashboard can therefore filter campaigns by:
-
Project
-
Activity
This allows campaigns belonging to different projects or activities to be monitored from the same Grist document.
| Field | Description |
|---|---|
| dataset_id | Dataset identifier |
| campaign | Associated campaign |
| publication_status | Current publication status |
| last_attempt_status | Status of the last publication attempt |
| finished_at | Timestamp of the last attempt |
| log_file | Associated log file |
Typical publication statuses are:
SUCCESS
FAILED
PENDING
| Field | Description |
|---|---|
| dataset_id | Dataset identifier |
| campaign | Associated campaign |
| run_id | Publication run identifier |
| started_at | Start timestamp |
| finished_at | Finish timestamp |
| status | Publication status |
| exit_code | Publisher exit code |
| log_file | Associated log file |
| error_message | Error information |
Create a Grist table with table ID Diagnostics and these columns before using
the default diagnostic synchronization:
| Field | Description |
|---|---|
| diagnostic_id | Unique diagnostic attempt identifier |
| diagnostic_run_id | Identifier shared by one diagnostic run |
| dataset_id | Dataset identifier |
| campaign | Associated campaign |
| started_at | Attempt start timestamp |
| finished_at | Attempt finish timestamp |
| outcome | RECOVERED, DIAGNOSED, UNCLASSIFIED, or EXECUTION_ERROR |
| publisher_status | Recognized publisher status, when present |
| exit_code | Publisher process exit code |
| http_status | Transaction service HTTP status |
| error_type | RFC problem type returned by the service |
| schema_url | STAC schema implicated by validation |
| rejected_value | Value rejected by schema validation |
| suggested_value | Closest accepted enum value, when identifiable |
| summary | Condensed diagnostic message |
| server_instance | Server-side problem instance identifier |
| log_file | Full local verbose log path |
| stac_file | Retained failed-item path, or requested per-run item path |
Note: Grist credentials are supplied via environment variables and are not stored in the repository.
The Grist document contains a custom dashboard widget providing:
-
Overall dataset KPIs
-
Published/failed/pending counts
-
Campaign-level publication progress
-
Project filtering
-
Activity filtering
-
Campaign progress bars
The dashboard is intended as an operational view of the DuckDB workflow state rather than as a replacement for the database.
Archival is separated from publication.
The publisher VM may lack write access to the final storage location. Instead, Pubflow generates a portable list of archive operations.
A dataset becomes eligible for archival when:
publication_status = SUCCESS
archive_status = PENDING
pubflow archive generate tipmip-cnrm
Use a limit for testing:
pubflow archive generate tipmip-cnrm --limit 10
The command generates a CSV containing:
| Column | Description |
|---|---|
| dataset_id | Dataset identifier |
| mapfile | Path to the mapfile |
| archive_path | Destination archive path |
Example:
dataset_id,mapfile,archive_path
CMIP6Plus.TIPMIP.CNRM-CERFACS.CNRM-ESM2-1.esm-piControl.r1i1p2f2.AERmon.cdnc.gr.v20231218,/modfs/esgf/topublish/CNRM-CERFACS/.mapfiles/....map,/mnt/scality/WCRP/CMIP6Plus/TIPMIP/CNRM-CERFACS/CNRM-ESM2-1/esm-piControl/.mapfiles/....map
Note: The archive path is dynamically generated from the dataset's DRS and the campaign's configured archive depth.
The archive.depth setting determines how deep the archive destination should be created within the project's DRS
hierarchy.
For example:
archive:
depth: experiment_id
can produce:
/mnt/scality/WCRP/CMIP6Plus/TIPMIP/CNRM-CERFACS/
CNRM-ESM2-1/
esm-piControl/
.mapfiles/
The DRS is interpreted using esgvoc.
The workflow does not hard-code project-specific DRS fields, allowing different projects to use their own vocabulary and DRS definitions.
The workflow uses esgvoc for DRS interpretation.
For example, for the cmip6plus project, ESGVOC provides DRS components including:
-
mip_era -
activity_id -
institution_id -
source_id -
experiment_id -
member_id -
table_id -
variable_id -
grid_label -
version
Note: Pubflow does not contain a CMIP6Plus-specific hard-coded
parse_drs()implementation. This allows different projects to use their own vocabulary and DRS definitions.
The generated archive task CSV is designed to be portable.
It can be transferred to a computing centre and executed independently:
python bin/archive.py archive_tasks.csv
The archive executor does not require:
-
DuckDB
-
Grist credentials
-
Publisher credentials
-
Access to the publisher VM
-
The Pubflow workflow database
It reads the task CSV and copies each mapfile to its specified archive destination.
Request a separate results CSV with:
python bin/archive.py archive_tasks.csv \
--results archive_results.csv
Result format:
| Column | Description |
|---|---|
| dataset_id | Dataset identifier |
| mapfile | Path to the mapfile |
| archive_path | Destination archive path |
| status | Status of the operation |
| error_message | Error message, if any |
Example:
TEST.DATASET,/source/test.map,/archive/.mapfiles/test.map,SUCCESS,
The current executor uses Python's shutil.copy2() and creates the destination directory when necessary.
esgf-publisher-workflow/
│
├── config/
│ ├── campaigns.yml
│ └── publisher.yml
│
├── db/
│ └── schema.sql
│
├── logs/
│
├── pubflow/
│ └── cli.py
│
├── workflow/
│ ├── archive.py
│ ├── campaign.py
│ ├── config.py
│ ├── database.py
│ ├── exporter.py
│ ├── executor.py
│ ├── grist.py
│ └── registry.py
│
├── bin/
│ ├── init_db.py
│ ├── publisher.py
│ └── archive.py
│
└── pyproject.toml
Note: The exact contents may evolve as the workflow develops.
External credentials and connection information are supplied via environment variables.
Repository-owned paths are resolved from the installed project location, so commands do not depend on the current working directory. Deployments can override them when required:
export PUBFLOW_DB_PATH=/path/to/publications.duckdb
export PUBFLOW_CONFIG_DIR=/path/to/config
export PUBFLOW_CAMPAIGNS_FILE=/path/to/campaigns.yml
Relative logging and campaign mapfile paths are resolved from the project
root. User-relative paths such as ~/.esg/esg.yaml.EASTINT and environment
variables in configured paths are expanded automatically.
export GRIST_BASE_URL=...
export GRIST_API_KEY=...
export GRIST_DOC_ID=...
These variables should not be committed to the repository.
For permanent local configuration, they can be added to the appropriate shell configuration, such as:
~/.bashrc
Pubflow intentionally separates responsibilities between the workflow manager, ESG publisher, and computing centre.
Pubflow is responsible for:
-
Campaign configuration
-
Dataset registration
-
Publication orchestration
-
Batch management
-
Retry handling
-
Publication state tracking
-
Logging
-
Status export
-
Grist synchronisation
-
Archive task generation
-
ESG publisher profile selection
The ESG publisher remains responsible for:
-
Dataset extraction
-
ESG metadata generation
-
ESG publication
-
STAC interaction
-
EGI Check-in authentication
-
Token management
The computing centre is responsible for:
-
Executing archive tasks
-
Writing mapfiles to the final archive location
-
Returning archive results
This separation keeps Pubflow lightweight and avoids duplicating functionality already provided by the ESG publisher.
-
Campaign configuration
-
Project/activity metadata
-
DuckDB database
-
Dataset registration
-
Publication workflow
-
Batch publication
-
Publication retries
-
Dry-run configuration
-
Publication logging
-
Validation workflow
-
CSV export
-
Grist synchronisation
-
Grist operational dashboard
-
Project/activity filtering in Grist
-
pubflowCLI -
ESGVOC-based DRS parsing
-
Configurable archive depth
-
Portable archive task generation
-
Standalone archive executor
-
Archive result CSV
-
Multiple ESG publisher configuration profiles
-
Runtime
--configselection foresgpublish -
Publisher-side mapfile path mappings
Potential future improvements include:
-
Safe handling of existing archive destination files
-
Resumable archive operations
-
Explicit archive conflict detection
-
Import of archive results into the workflow database
-
More detailed Grist campaign/run visualisations
-
Additional operational monitoring for long-running publication processes
-
Optional inspection of selected ESG publisher configuration values
Authentication/token management is not currently planned for Pubflow, as this functionality is already handled by
esgpublish.
A standard publication campaign follows these steps:
# Initialize the database
python bin/init_db.py
# Load campaign definitions
python workflow/campaign.py
# Register datasets
pubflow dataset register tipmip-cnrm
# Optionally validate datasets
pubflow dataset validate tipmip-cnrm
# Test a small publication batch
pubflow publication run tipmip-cnrm --limit 10
# Publish the campaign
pubflow publication run tipmip-cnrm
# Export status
pubflow report export --output campaign_status.csv
# Synchronize status to Grist
pubflow grist sync
# Generate archive tasks
pubflow archive generate tipmip-cnrm
# Transfer archive_tasks.csv to the computing centre
# Execute archive tasks there
python bin/archive.py archive_tasks.csv \
--results archive_results.csv
For long-running publication campaigns, running Pubflow inside a persistent terminal session such as tmux is
recommended.