Repository navigation
[bug] Fix Flink TaskManager Metaspace retention via shared Hadoop ReflectionUtils cache - #9827
zhang-arvin wants to merge 1 commit into
Conversation
…pache#9795) Parquet 1.15.x CodecFactory instantiates codecs through Hadoop's ReflectionUtils.CONSTRUCTOR_CACHE, a static ConcurrentHashMap with strong keys living in the shared AppClassLoader. Paimon shades parquet but not Hadoop, so the shaded ZstandardCodec class loaded by each job's ChildFirstClassLoader is pinned there forever (~18 MB Metaspace per finished Flink job, never reclaimed). Parquet 1.16.0 rewrites CodecFactory to construct codecs directly (DirectCodecFactory) and no longer routes through the shared ReflectionUtils cache, so the job classloader is reclaimable.
JingsongLi
left a comment
There was a problem hiding this comment.
Requirement fit: UNSUPPORTED for this version-only patch. Implementation: FINDINGS.
Reviewed be96f81ee6bb. The classloader-retention problem in #9795 is valuable, but changing only the dependency version does not remove its retaining reference. I am closing this version-only proposal because it does not deliver the stated fix and broadens a maintenance-branch dependency without solving the reported failure. Please replace the codec-instantiation path that populates the shared Hadoop cache and add a regression demonstrating that the job-loaded codec class is no longer retained; that would be a useful focused follow-up or basis for reopening.
Validation: inspected the published 1.16.0 CodecFactory source and bytecode, then ran a small Java probe using that artifact. The Zstandard codec key was absent before construction, present after getDecompressor(ZSTD), and still present after release(). This verifies the cache insertion; it is not a full Flink heap-dump reproduction. CI also currently has failed builds.
| <iceberg.version>1.6.1</iceberg.version> | ||
| <hudi.version>0.15.0</hudi.version> | ||
| <parquet.version>1.15.2</parquet.version> | ||
| <parquet.version>1.16.0</parquet.version> |
There was a problem hiding this comment.
[P1] Remove the shared-cache call instead of only upgrading Parquet
Parquet 1.16.0 still calls ReflectionUtils.newInstance(codecClass, ...) in CodecFactory.getCodec. Its constructor does not switch to DirectCodecFactory; that is a separate factory method. Paimon's vendored ParquetWriter also still constructs new CodecFactory(...). A probe against 1.16.0 leaves ZstandardCodec in ReflectionUtils.CONSTRUCTOR_CACHE after release(), so a shared Hadoop loader continues to retain the job-loaded codec class in the reported session-cluster scenario. Please fix that instantiation path and verify that the shared cache never gains the job class.
Purpose
Fixes #9795
Root cause
Every finished Flink batch job shipping the Paimon connector as a user jar leaks ~18.2 MB TaskManager Metaspace permanently. The retained object is the job's
ChildFirstClassLoaderclass fororg.apache.paimon.shade.org.apache.parquet.hadoop.codec.ZstandardCodec, pinned by Hadoop'sReflectionUtils.CONSTRUCTOR_CACHE(static, strong keys) in the sharedAppClassLoader.Paimon shades parquet (1.15.2 on release-1.3) but deliberately does not shade Hadoop, so parquet's
CodecFactoryhands the job-classloader-loaded codec class into the shared Hadoop utility that never evicts production-side.Fix
Parquet 1.16.0 rewrites
CodecFactoryto instantiate codecs directly (DirectCodecFactory) instead of going throughReflectionUtils.newInstance, so the strong-keyed shared cache is no longer on the path. The job classloader becomes garbage-collectible after job completion. main/2.1-SNAPSHOT already runs 1.16.0; this backports the version to release-1.3.Verification
CodecFactory: invokesReflectionUtils.newInstance→ shared cache. 1.16.0CodecFactoryconstructor buildsDirectCodecFactorydirectly — no ReflectionUtils reference.Reporter's verification method (from issue thread)
Heap-dump count of user-classloader keys in
CONSTRUCTOR_CACHEshould drop to 0 after upgrade.