Skip to content

[bug] Fix Flink TaskManager Metaspace retention via shared Hadoop ReflectionUtils cache - #9827

Closed
zhang-arvin wants to merge 1 commit into
apache:release-1.3from
zhang-arvin:fix/issue-9795
Closed

zhang-arvin wants to merge 1 commit into
apache:release-1.3from
zhang-arvin:fix/issue-9795

Conversation

@zhang-arvin

Copy link
Copy Markdown
Contributor

Purpose

Fixes #9795

Root cause

Every finished Flink batch job shipping the Paimon connector as a user jar leaks ~18.2 MB TaskManager Metaspace permanently. The retained object is the job's ChildFirstClassLoader class for org.apache.paimon.shade.org.apache.parquet.hadoop.codec.ZstandardCodec, pinned by Hadoop's ReflectionUtils.CONSTRUCTOR_CACHE (static, strong keys) in the shared AppClassLoader.

Paimon shades parquet (1.15.2 on release-1.3) but deliberately does not shade Hadoop, so parquet's CodecFactory hands the job-classloader-loaded codec class into the shared Hadoop utility that never evicts production-side.

Fix

Parquet 1.16.0 rewrites CodecFactory to instantiate codecs directly (DirectCodecFactory) instead of going through ReflectionUtils.newInstance, so the strong-keyed shared cache is no longer on the path. The job classloader becomes garbage-collectible after job completion. main/2.1-SNAPSHOT already runs 1.16.0; this backports the version to release-1.3.

Verification

  • Bytecode check of 1.15.2 CodecFactory: invokes ReflectionUtils.newInstance → shared cache. 1.16.0 CodecFactory constructor builds DirectCodecFactory directly — no ReflectionUtils reference.
  • Link check: same vendored parquet sources + shade config compile on main with 1.16.0 (already proven upstream).

Reporter's verification method (from issue thread)

Heap-dump count of user-classloader keys in CONSTRUCTOR_CACHE should drop to 0 after upgrade.

…pache#9795)

Parquet 1.15.x CodecFactory instantiates codecs through Hadoop's
ReflectionUtils.CONSTRUCTOR_CACHE, a static ConcurrentHashMap with
strong keys living in the shared AppClassLoader. Paimon shades parquet
but not Hadoop, so the shaded ZstandardCodec class loaded by each job's
ChildFirstClassLoader is pinned there forever (~18 MB Metaspace per
finished Flink job, never reclaimed).

Parquet 1.16.0 rewrites CodecFactory to construct codecs directly
(DirectCodecFactory) and no longer routes through the shared
ReflectionUtils cache, so the job classloader is reclaimable.

@JingsongLi JingsongLi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requirement fit: UNSUPPORTED for this version-only patch. Implementation: FINDINGS.

Reviewed be96f81ee6bb. The classloader-retention problem in #9795 is valuable, but changing only the dependency version does not remove its retaining reference. I am closing this version-only proposal because it does not deliver the stated fix and broadens a maintenance-branch dependency without solving the reported failure. Please replace the codec-instantiation path that populates the shared Hadoop cache and add a regression demonstrating that the job-loaded codec class is no longer retained; that would be a useful focused follow-up or basis for reopening.

Validation: inspected the published 1.16.0 CodecFactory source and bytecode, then ran a small Java probe using that artifact. The Zstandard codec key was absent before construction, present after getDecompressor(ZSTD), and still present after release(). This verifies the cache insertion; it is not a full Flink heap-dump reproduction. CI also currently has failed builds.

Comment thread pom.xml
<iceberg.version>1.6.1</iceberg.version>
<hudi.version>0.15.0</hudi.version>
<parquet.version>1.15.2</parquet.version>
<parquet.version>1.16.0</parquet.version>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Remove the shared-cache call instead of only upgrading Parquet

Parquet 1.16.0 still calls ReflectionUtils.newInstance(codecClass, ...) in CodecFactory.getCodec. Its constructor does not switch to DirectCodecFactory; that is a separate factory method. Paimon's vendored ParquetWriter also still constructs new CodecFactory(...). A probe against 1.16.0 leaves ZstandardCodec in ReflectionUtils.CONSTRUCTOR_CACHE after release(), so a shared Hadoop loader continues to retain the job-loaded codec class in the reported session-cluster scenario. Please fix that instantiation path and verify that the shared cache never gains the job class.

@JingsongLi JingsongLi closed this Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants