Skip to content

Share SQL/JSON across Flink lines and preserve legacy ENCODE bytes - #259

Merged
jordepic merged 2 commits into
mainfrom
feat/flink118-encode-json
Sep 26, 2026
Merged

jordepic merged 2 commits into
mainfrom
feat/flink118-encode-json

Conversation

@jordepic

Copy link
Copy Markdown
Collaborator

Flink 1.18 SQL declares ENCODE as BINARY(1) while returning variable-length bytes. The existing batch evaluator now preserves the full values using Arrow Binary internally, without changing the public schema or Flink expressions. Final projections and fused consumers are admitted; incompatible operator/native-sink boundaries retain explicit fallback. Casting to BYTES establishes an ordinary variable-width boundary.

SQL/JSON now uses one native reader across 1.18 and 2.2.1. Small compatibility adapters select the released Jackson recycler contract, floating-number behavior and parsing limits. The 1.18 profile requires Jackson 2.14.2/JDK 17 and preserves Double formatting, signed zero, overflow and unrestricted nesting; 2.2 retains its Jackson 2.18.2/BigDecimal behavior. Buffer-history and double-bit JNI tests are shared between profiles. Unsupported runtimes retain the host evaluator.

Validation uses released Flink 1.18.1 and 2.2.1 on JDK 17 with the debug native build:

  • 1.18: 735 distinct targeted cases, 714 passed and 21 version-specific skips.
  • 2.2.1: 727 targeted cases passed.
  • 106 native function tests passed.
  • Coverage includes charset errors and evaluation order, public schemas, Parquet boundaries, shared binary UDFs, randomized numeric values, recycler history and 4,096-level JSON nesting.

Coverage and fallback documentation is updated. No throughput claim or full upstream inventory refresh is included. This PR is stacked on #258; Delta remains outside its scope.

root added 2 commits September 26, 2026 13:12
Flink 1.18 SQL declares BINARY(1) for ENCODE while returning variable-length byte arrays. Reuse the shared batch evaluator and carry affected scalar results as Arrow Binary internally, keeping the original expressions and public schema intact.

Keep explicit fallbacks where fixed-width metadata would cross another operator or a native sink. A cast to BYTES establishes the ordinary variable-width boundary. Document the admitted forms and retained limits.

Validated 31 ENCODE/DECODE/charset and boundary cases on Flink 1.18.1, the 22 shared charset cases on 2.2.1, and 45 shared binary-UDF/lifecycle regressions on each line. This adds coverage through the existing JVM batch bridge; no throughput improvement is claimed.
Flink 1.18 uses Jackson 2.14.2 with Double JSON values and no newer parsing limits, while 2.2 uses Jackson 2.18.2 with BigDecimal values. Select those semantics through small verified runtime adapters while sharing parsing, SIMD selection, path handling, and buffer ownership.

Reuse the existing JDK 17 double formatter and grow the legacy parser stack for deep nesting. Keep unverified runtimes on the batch host evaluator. Move recycler-history and double-bit parity fixtures into the common suite and document the supported profiles.

Validation: 735 distinct targeted cases on Flink 1.18.1 (714 passed, 21 version-specific skips), 727 passed on Flink 2.2.1, and 106 native function tests passed. Direct JNI checks cover randomized doubles, overflow, signed zero, extreme exponents, buffer history, and 4096-level nesting. These are correctness and coverage checks with a debug native build, not throughput measurements.
@jordepic
jordepic changed the base branch from feat/flink118-udf-lookup-int96 to main September 26, 2026 20:26
@jordepic
jordepic marked this pull request as ready for review September 26, 2026 20:26
@jordepic
jordepic merged commit 022f735 into main Sep 26, 2026
47 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant