Skip to content

Close Flink 1.18 legacy UDF, lookup and INT96 coverage gaps - #258

Merged
jordepic merged 4 commits into
mainfrom
feat/flink118-udf-lookup-int96
Sep 26, 2026
Merged

jordepic merged 4 commits into
mainfrom
feat/flink118-udf-lookup-int96

Conversation

@jordepic

Copy link
Copy Markdown
Collaborator

Flink 1.18 now routes deprecated scalar registrations and legacy sync/async lookup sources through the existing native operators, and Parquet sinks support the host’s default INT96 timestamps. This closes the three selected 1.18 parity gaps while preserving the existing UDF specialization, changelog and local-time INT64 fallbacks.

Stacked on #257 (fix/flink118-stock-rocksdb); the diff excludes its backend substitution work.

  • Legacy UDFs retain constructor state, job parameters, lifecycle and shared-stateful-call guards through the host-version adapter.
  • Legacy lookup sources use Flink’s released provider resolution and generated runners with the existing Arrow batch boundaries. Both legacy and modern upstream contracts require native row processing.
  • INT96 uses released parquet-rs column writers beside ordinary Arrow encoders, with nested definition/repetition levels and byte-bounded row groups. A batched JVM callback preserves Flink’s local calendar/DST conversion. Flink retains filesystem and checkpoint ownership. The unused timestamp unit is ignored in INT96 mode, matching the host.

Validation:

  • Focused Java regressions: 129 passed, 2 skipped on 1.18; all 124 passed on a clean 2.2 build.
  • Native Parquet: all 26 unit tests passed. Eight physical-file comparisons on each host line cover extreme dates, local zones, nested null/empty collections, sliced batches and multiple row groups.
  • Unchanged upstream 1.18: 262 UDF/lookup invocations passed; all 26 execution contracts proved native row processing, including legacy variants.
  • Unchanged upstream 1.18 Parquet: all 8 cases passed. All four timestamp sink plans were admitted without fallback; the timestamp test class created a native writer.
  • After the ignored-unit correction, all 30 focused planner/SQL sink tests passed on 2.2 and the upstream Parquet suite passed again with SQL inventory enabled.

These are scoped validations, not a refreshed full-suite acceleration count. No release benchmark or throughput improvement is claimed. Shared coverage gaps and Delta are outside this change.

root added 4 commits September 26, 2026 11:55
Resolve legacy and modern scalar definitions through the host-version adapter so deprecated registrations retain the existing columnar evaluator, constructor state, lifecycle and job parameters. Preserve specialization and shared mutable-call guards instead of creating a separate legacy evaluator.

Add legacy registration parity and fallback tests plus native execution contracts for the unchanged upstream rich-function cases. Focused UDF regressions pass on Flink 1.18 and 2.2. This extends coverage within native islands; no standalone throughput improvement is claimed.
Use the released host provider resolution and generated runners for legacy lookup sources, keeping the same Arrow probe and result boundaries as modern sources. Preserve synchronous and asynchronous lifecycle, constant keys, residual filtering and retry semantics without introducing a second lookup runtime.

Require native row-processing evidence for legacy upstream lookup contracts. New sync/async fixtures cover duplicate matches, misses, nulls, dimension filters and lifecycle; existing modern lookup regressions pass on both supported host lines. This is coverage parity, without a new throughput claim.
Keep standard Arrow column encoders for ordinary leaves and use typed INT96 encoders for timestamp leaves within the same row group. Preserve nested definition and repetition levels and account for buffered timestamp values in row-group sizing. Flink continues to own filesystem writes and checkpoint commits.

UTC encoding preserves the host millisecond/fraction arithmetic without narrowing to epoch nanoseconds. Local INT96 uses one JVM callback per timestamp column to preserve the host calendar, default timezone and DST behavior. Local-time INT64 remains a planner fallback.

Validation: 26 native Parquet unit tests; eight physical-file timestamp comparisons on each host line covering extreme dates, local zones, nested null/empty collections, slices and multiple row groups; partitioned SQL sink parity and admission tests. Combined focused Java regressions: 129 passed and 2 skipped on 1.18, all 124 passed on a clean 2.2 build. Coverage extension only; no release benchmark or throughput claim.
The unchanged 1.18 timestamp fixture supplies an unused time-unit value in INT96 mode. Flink ignores it, but native admission still validated it, leaving a silent fallback after INT96 encoding was added. Validate units only for INT64 and pin ignored-unit behavior through host/native SQL sink parity.

Validation: all 30 focused planner and SQL sink tests pass on 2.2. All eight unchanged upstream 1.18 Parquet cases pass; all four timestamp sink plans are admitted without fallback and the timestamp class creates a native writer. Also record the successful 262-case upstream UDF/lookup run and its 26 native execution contracts.
@jordepic
jordepic marked this pull request as ready for review September 26, 2026 20:25
@jordepic
jordepic changed the base branch from fix/flink118-stock-rocksdb to main September 26, 2026 20:25
@jordepic
jordepic merged commit 91463b5 into main Sep 26, 2026
58 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant