Skip to content

Detect calibration damage between any two images: signed per-source statistics and local noise #90

Description

@Athanaseus

Scope

This is about comparing any two images of the same field, not specifically DD vs DI. The question is the general one, aimfast already exists to answer: did this processing step improve things or make them worse? — for example:

  • successive self-calibration rounds (did round N+1 actually help?)
  • before/after a flagging strategy
  • two calibration approaches, two pipelines, two parameter choices
  • direction-independent vs direction-dependent calibration (the test case below)

aimfast already supports pairwise image and catalogue comparison via --compare-images and --compare-models, including flux-suppression checks. What is missing is the ability to detect localised, signed damage, which is invisible to every statistic currently reported.

The DI/DD pair below is simply a case where the damage is large, well understood, and independently confirmed, which makes it a good test case for the metrics proposed here.

Summary

An image can lose 8–11% of its source catalogue while every statistic aimfast currently reports says it is equivalent to its counterpart. This proposes the measurements that would catch it.

The test case

DI (killMS/DDFacet) vs DD (CubiCal) restored images of the same field, both DDFacet 0.5.3.1, 6075² @ 1.5″, same phase centre. The DD run peeled a bright source and over-subtracted it.

What aimfast reports today:

statistic DI DD verdict
global MAD rms 9.4 µJy 9.4 µJy identical
deepest_negative (residual, peak box) 83.5 93.4 DD better
matched-source flux ratio — 1.00–1.05 no suppression

What actually happened:

DI DD
brightest source +162.24 mJy −24.74 mJy (over-subtracted hole)
sources (aegean, 5σ, island level) 3905 3482 (−10.8%)
sources (pybdsf) 3856 3537 (−8.3%)
sources (breizorro) 4163 3695 (−11.2%)

Three independent finders at matched thresholds and matched units. The surviving sources keep their correct fluxes — the damage is entirely source loss, not dimming.

Mechanism — and why global statistics cannot see it

estimator kind DI DD change
MAD rms global 9.4 µJy 9.4 µJy 0%
breizorro noise out median of local map 8.37 µJy 8.33 µJy −0.5%
aegean local_rms local 9.59 µJy 10.47 µJy +9.1%
pybdsf Isl_rms local 10.34 µJy 11.09 µJy +7.2%
NB: breizorro builds a genuinely local noise map (minimum filter over 50×50 boxes) but then floors it at its own median, so quiet regions are clamped to one value and only noisier-than-median regions keep local detail. The number quoted here is the median.

Every global estimator calls the images identical; every local estimator finds DD 7–9% worse. The artefacts are structured and localised, so averaging over millions of empty pixels hides them.

Note that the loss is not purely a raised-threshold effect: breizorro used the same absolute cutoff for both images (0.04 mJy, based on its global noise) and still lost 468 sources. So negative holes and corrupted islands directly interfere with detection and raise the local floor.

Control: Is this just a finder disagreement?

Finder-vs-finder cross-matches on the same image, versus DI-vs-DD with the same finder:

match rate asymmetry (A-only ÷ B-only)
control — 2 finders, same image (5 pairs) 83.1 – 88.1% 0.64 – 1.08
signal — 1 finder, DI vs DD (3 finders) 74.7 – 78.7% 1.48 – 2.04

No overlap on either axis. Finders disagree roughly symmetrically (each finds what the other missed); DI-vs-DD loss is one-directional. And inter-finder agreement is identical on both images (83.3% / 83.3%), so DD did not make detection unreliable; the lost sources are lost consistently by all finders, i.e. real sources falling below threshold.

What to add

1. Signed statistics per source box (main change)

_source_residual_results already does most of the work where it loops over skymodel sources, cuts a box around each, compares two images, and records:

[res1_rms, res2_rms, res1_rms/res2_rms, phase_centre_dist,  model_source.name, model_source.flux.I]

but the only statistic per box is std(), which is sign-blind: a −24 mJy hole and a +24 mJy positive residual raise it identically.

Proposed: also record min() and SUM_NEG per box. Both are already defined in residual_image_stats, so no new metric definitions are needed. Two useful breakdowns then come free from fields already stored:

  • model_source.flux.I → suppression vs source brightness
  • phase_centre_dist → suppression vs off-axis distance, which discriminates primary-beam error (smooth, radial about the pointing centre) from DD-solve error (discontinuous at facet edges)

2. Report local noise, not only global

residual_image_stats returns only global statistics (MEAN/STDDev/RMS/MAD/…), which is why this run looks equivalent. The source finders already compute a local rms map and write it per source (aegean local_rms, pybdsf Isl_rms), so the quantity is available at no cost.

3. The population statistic should be a count, not a min or a median

Prototyped on 300 common positions, measuring each source's local minimum ÷ its own peak:

percentile        DI         DD
    50%       -0.038     -0.049      median: nearly identical
     1%       -0.129     -0.260      tail: 2x divergence
  worst       -0.427     -0.406      worst: DD marginally BETTER

holes deeper than 10% of own peak:  DI 6/298    DD 17/298   (2.8x)

Two non-obvious results:

  • the median does not discriminate — a median-based metric passes this run - the single worst value does not discriminate either — DI has its own one bad source, so "deepest negative in the field" fails

The tail population discriminates: fraction of sources carrying a hole deeper than X% of their own peak. So the metric should be a counting statistic over the source population.

Note that DI is not clean either (6/298 with >10% holes) because real deconvolution leaves negative values near bright sources. The claim is relative, not absolute.

Why this matters

Currently, the only way to notice this failure is to eyeball images and cross-match catalogues by hand. An image that passes every existing check can still have lost 11% of its faint sources, which directly corrupts source counts, completeness, and anything derived from them; at the survey scale, hand inspection of every pointing is not possible.

The same measurements answer the everyday question directly: Did this self-cal round help?
A round that lowers the global rms while quietly deepening negatives around sources, or losing faint detections, currently looks like an improvement. Signed per-source statistics and a local noise comparison would show it for what it is — and they apply to any image pair, not just a
DI/DD one.

Tests

  • Synthetic pair with a known negative hole at a known source position: assert the signed statistic flags it, and that std() alone does not.
  • Assert the flux and off-axis breakdowns are produced from the already-stored fields.
  • Assert the population metric is a count over sources, and that median/min variants are not used as the primary discriminator (regression guard for the result above).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions