Skip to content

Independent reproduction and overclaim checklist before paper text #16

Description

@esaran1

Why

The package can run, pass tests, and still not support the sentence in the paper. Someone who did not implement the audit should try to reproduce tables from frozen artifacts and reject overclaims.

This repo contains ensemble projection, racing, conformal intervals, and ASLA-Bench as software. Those names are real methods here. They are not measured-BPB evidence until the 45-job grid exists.

Checklist

  • Frozen runs.parquet SHA-256 matches audit_metadata.input_sha256
  • Command line matches preregistered estimand, budgets, intermediate holdout, and fit form
  • Projection vs single-scale is the primary comparison
  • Ranker intervals are described as seed-bootstrap percentile intervals, not conformal
  • Conformal projection intervals are named conformal only when that method was used
  • Gate is described as a one-step heuristic; racing is a separate cost-aware policy, not error control for the gate
  • Saturation fit is blind is kept as a limitation
  • Under-seeded cells and unestimated noise bands are reported
  • No mechanism-crossover numbers
  • No results from data/example_runs.csv or --fast
  • Demo stdout is not the source of paper tables

What to do

Have a second person clone the frozen SHA, rerun asla audit and the table generator (once that script exists), and confirm bit-identical JSON/tables or document allowed nondeterminism.

Do not fabricate

  • Soften language to hide overlapping intervals or a blind fit
  • Present synthetic ASLA-Bench sweeps as the 45-job C4-EN result

Metadata

Metadata

Assignees

No one assigned

    Labels

    paperPaper tables, figures, and write-upscienceScientific design, estimands, and claim language

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions