Goal
Extend the v9 unification (#215) by collapsing the 6 *OnDieFabric classes (CPU/GPU/NPU/CGRA/DPU/TPU) into an inheritance hierarchy: a shared OnDieFabric base in compute_block_common carries the 7 shared fields + optional mesh dims + optional confidence; per-block-kind subclasses contribute only the architecture-specific topology enum.
Reconciles the endpoint-count naming inconsistency at the same time: CPU's stop_count and GPU's controller_count are renamed to unit_count (matches the existing convention used by NPU/CGRA/DPU/TPU; per-arch field description documents what a ""unit"" means in each context).
Why now
After the v9 sprint (#215) closed, compute_block_common hosts 4 re-exported primitives + 2 unified types (TheoreticalPerformance + ThermalProfile). The v9 paper exercise flagged OnDieFabric as the v10 candidate:
OnDieFabric unification: 3 of 6 classes (NPU/CGRA/DPU) are byte-identical (10 fields). TPU drops mesh dims; CPU uses stop_count; GPU uses controller_count for the same conceptual ""endpoint count"" field. Naming reconciliation + extension model needed. Defer to v10.
Why inheritance (vs alias)
Unlike v8/v9 where the 5 (resp. 5) target classes were byte-identical and collapsed via aliases, OnDieFabric is NOT byte-identical across the 6 block kinds because each carries its own topology: <Arch>NoCTopology enum. The topology enums are genuinely architecture-specific (e.g., CPU's infinity_fabric is meaningless to a TPU; DPU's aie_mesh is meaningless to a CPU). Preserving them is correct.
The inheritance approach gives the best of both worlds:
- Shared fields defined once (in base)
- Per-block-kind enum validation preserved
isinstance(x, NPUOnDieFabric) AND isinstance(x, OnDieFabric) both work
- Each per-kind subclass shrinks from ~30 LOC to ~5 LOC
Survey: 6 OnDieFabric classes today
| Class |
Shared 7 fields |
mesh_cols/rows |
confidence |
endpoint name |
| NPUOnDieFabric |
YES |
YES |
YES |
unit_count |
| CGRAOnDieFabric |
YES |
YES |
YES |
unit_count |
| DPUOnDieFabric |
YES |
YES |
YES |
unit_count |
| TPUOnDieFabric |
YES |
no |
YES |
unit_count |
| CPUOnDieFabric |
YES |
no |
no |
stop_count |
| GPUOnDieFabric |
YES |
no |
no |
controller_count |
Shared 7 fields (all 6 have these):
bisection_bandwidth_gbps, flit_size_bytes, hop_latency_ns, pj_per_flit_per_hop, routing_distance_factor, topology (per-kind enum), endpoint-count-field.
Topology enums (stay separate)
Each block kind has its own topology enum with genuinely different values:
- CPU: ring / double_ring / io_die_plus_ccd / infinity_fabric
- GPU: crossbar / ring / hierarchical
- NPU: dataflow_ring / systolic / crossbar / shared / partitioned
- CGRA: crossbar
- DPU: aie_mesh / crossbar
- TPU: crossbar / multi_crossbar
Only ""crossbar"" appears in multiple enums. Unifying into a single enum would lose architectural specificity. Topology enums stay per-block-kind.
YAML impact
3 YAMLs reference the to-be-renamed fields:
intel/intel_core_i7_12700k.yaml (CPU stop_count -> unit_count)
nvidia/jetson_agx_orin_64gb.yaml (GPU controller_count -> unit_count)
nvidia/jetson_agx_thor_128gb.yaml (GPU controller_count -> unit_count)
Same migration as any field rename: PR 3 updates these 3 YAMLs atomically with the schema change.
Sprint shape (3 PRs, mirror of v8 unification + v9 unification pattern)
-
PR 1 (graphs docs) -- paper exercise at docs/designs/v10-on-die-fabric-unification.md. Audits all 6 *OnDieFabric classes side-by-side, justifies the inheritance approach (vs alias), documents the endpoint-count rename + YAML migration plan.
-
PR 2 (embodied-schemas) -- add OnDieFabric base class to compute_block_common.py with 7 shared fields + optional mesh_cols / mesh_rows / confidence. Additive only: existing per-kind classes unchanged.
-
PR 3 (embodied-schemas) -- migrate all 6 per-kind classes to inherit from OnDieFabric base. CPU and GPU also get the stop_count -> unit_count and controller_count -> unit_count rename + 3 YAML updates in the same commit.
Expected diff: net -150 to -200 LOC of duplicated field definitions removed; 3 YAML field renames.
Out of scope (defer to v11)
has_external_dram vs has_host_dram naming reconciliation (touches more SKU YAMLs; semantic difference matters -- chip-attached HBM vs host-bus PCIe DRAM)
KPUNoCSpec (oldest module; pre-pattern; doubly-purposed)
- Topology enum unification (intentionally architecture-specific)
Refs
Goal
Extend the v9 unification (#215) by collapsing the 6
*OnDieFabricclasses (CPU/GPU/NPU/CGRA/DPU/TPU) into an inheritance hierarchy: a sharedOnDieFabricbase incompute_block_commoncarries the 7 shared fields + optional mesh dims + optional confidence; per-block-kind subclasses contribute only the architecture-specifictopologyenum.Reconciles the endpoint-count naming inconsistency at the same time: CPU's
stop_countand GPU'scontroller_countare renamed tounit_count(matches the existing convention used by NPU/CGRA/DPU/TPU; per-arch field description documents what a ""unit"" means in each context).Why now
After the v9 sprint (#215) closed,
compute_block_commonhosts 4 re-exported primitives + 2 unified types (TheoreticalPerformance+ThermalProfile). The v9 paper exercise flaggedOnDieFabricas the v10 candidate:Why inheritance (vs alias)
Unlike v8/v9 where the 5 (resp. 5) target classes were byte-identical and collapsed via aliases,
OnDieFabricis NOT byte-identical across the 6 block kinds because each carries its owntopology: <Arch>NoCTopologyenum. The topology enums are genuinely architecture-specific (e.g., CPU'sinfinity_fabricis meaningless to a TPU; DPU'saie_meshis meaningless to a CPU). Preserving them is correct.The inheritance approach gives the best of both worlds:
isinstance(x, NPUOnDieFabric)ANDisinstance(x, OnDieFabric)both workSurvey: 6 OnDieFabric classes today
Shared 7 fields (all 6 have these):
bisection_bandwidth_gbps,flit_size_bytes,hop_latency_ns,pj_per_flit_per_hop,routing_distance_factor,topology(per-kind enum), endpoint-count-field.Topology enums (stay separate)
Each block kind has its own topology enum with genuinely different values:
Only ""crossbar"" appears in multiple enums. Unifying into a single enum would lose architectural specificity. Topology enums stay per-block-kind.
YAML impact
3 YAMLs reference the to-be-renamed fields:
intel/intel_core_i7_12700k.yaml(CPU stop_count -> unit_count)nvidia/jetson_agx_orin_64gb.yaml(GPU controller_count -> unit_count)nvidia/jetson_agx_thor_128gb.yaml(GPU controller_count -> unit_count)Same migration as any field rename: PR 3 updates these 3 YAMLs atomically with the schema change.
Sprint shape (3 PRs, mirror of v8 unification + v9 unification pattern)
PR 1 (graphs docs) -- paper exercise at
docs/designs/v10-on-die-fabric-unification.md. Audits all 6*OnDieFabricclasses side-by-side, justifies the inheritance approach (vs alias), documents the endpoint-count rename + YAML migration plan.PR 2 (embodied-schemas) -- add
OnDieFabricbase class tocompute_block_common.pywith 7 shared fields + optionalmesh_cols/mesh_rows/confidence. Additive only: existing per-kind classes unchanged.PR 3 (embodied-schemas) -- migrate all 6 per-kind classes to inherit from
OnDieFabricbase. CPU and GPU also get thestop_count->unit_countandcontroller_count->unit_countrename + 3 YAML updates in the same commit.Expected diff: net -150 to -200 LOC of duplicated field definitions removed; 3 YAML field renames.
Out of scope (defer to v11)
has_external_dramvshas_host_dramnaming reconciliation (touches more SKU YAMLs; semantic difference matters -- chip-attached HBM vs host-bus PCIe DRAM)KPUNoCSpec(oldest module; pre-pattern; doubly-purposed)Refs
graphs#208(closed)graphs#215(closed)graphs/docs/designs/v9-thermal-profile-unification.md(risk TASK-2026-014: Phase 0.5 — PowerMeter abstraction + composition-test CI gate #4 / summary table)