Skip to content

Federated statistics histograms count infinite samples twice #5360

Description

@sylvesterkaczmarek

On main at a02743ebe1f7f42e601633685735596062fc7764, get_std_histogram_buckets adds infinite samples to the first/last bins and then appends another degenerate infinity bin containing samples already counted.

Reproduction

import numpy as np
from nvflare.app_common.abstract.statistics_spec import BinRange
from nvflare.app_common.statistics.numpy_utils import get_std_histogram_buckets

bins = get_std_histogram_buckets(
    np.array([-np.inf, 0., np.inf]), 3, BinRange(-1., 1.)
)
print(len(bins), sum(b.sample_count for b in bins))

Actual: 4 4, despite requesting three bins for three samples.
Expected: three bins and a total count of three, with negative and positive infinities included once in their respective edge buckets.

There is a related one-bin case: an elif prevents both infinity endpoints from being added to the same bucket. Using independent endpoint checks preserves a single bucket with bounds (-inf, inf) when both are present.

The duplicate counts also reach the dataframe statistics implementation and global histogram aggregation. Reproduced on macOS, Python 3.12.11 and NumPy 2.5.3, using an explicitly supplied finite histogram range. Automatic range inference with infinite values is outside this report.

Activity

  1. v0ropaev commented on Oct 7, 2026

    @v0ropaev
    Contributor

    Opened #5383 for this. Both parts, the appended bucket and the elif on the one-bin case.

    Worth adding to your writeup: when both infinities are present the negative bucket never makes it out either, since bucket is assigned twice in that block and the positive assignment overwrites it. So the output is three bins plus one (inf, inf), with the negative infinity counted once in bucket 0 and the positive one counted twice.

    The order dependence you mentioned is concrete in accumulate_hists: it keys the global histogram on the first client bins and then iterates only those when writing counts back, so a later client (inf, inf) bucket is dropped while the same bucket from the first client sticks and inflates the global total.

  2. sylvesterkaczmarek commented on Oct 7, 2026

    @sylvesterkaczmarek
    ContributorAuthor

    Thanks @v0ropaev — #5383 covers both parts of the histogram defect, including the one-bin endpoint case and the downstream order-dependence source. I re-reviewed and approved it.

  3. sylvesterkaczmarek commented on Oct 7, 2026

    @sylvesterkaczmarek
    ContributorAuthor

    Thanks @v0ropaev — #5383 covers both parts of the histogram defect, including the one-bin endpoint case and the downstream order-dependence source. I re-reviewed and approved it.

  4. v0ropaev commented on Oct 7, 2026

    @v0ropaev
    Contributor

    Correction to my comment above: I closed #5383. #5361 already does this and is the better patch, and I had missed it because I read this thread and never searched the open PRs. Go with #5361.

    The one observation from mine that is not in this thread: with both infinities present, the appended bucket for minus infinity never reached the output either, because bucket is assigned twice in that block and the positive assignment overwrote it. #5361 deletes the block, so it is moot, but it explains why the symptom looked like one extra bucket rather than two.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions