Repository navigation
feat(loader): upgrade Java 17 for Spark-API - #785
Merged
Merged
Conversation
- manage compiler, surefire and coverage plugin versions - keep Java 8 output and use compatible Mockito 4 - resolve optional server repositories, refs and pins - verify checkout and package servers without install - share authenticated startup and filter AppleDouble files
- connect repository and fetch refs to all server installers - reuse the shared installer for Tools and Spark validation - preserve Loader source caching with verified checkouts - isolate fixture builds and retain separate HTTPS RPC ports
- preserve Mockito 2 interaction assertions with the equivalent API - align test-only Byte Buddy dependencies with Mockito 4 - retain the Spring Boot 2 and Java 11 shared baseline
- run shared server selection and installer tests in one CI job - document Maven 3.6.3 as the minimum source build version - report missing or ambiguous server archives before exiting
- trigger push checks for client installer changes - trigger pull request checks for the same source path - cover the installer exercised by the existing contracts
- encode legacy graph configuration as properties - keep task and target tests aligned with server contracts - inherit shared compiler and coverage plugin versions - preserve Java 8 APIs in compatibility tests
- align Spark 3.5.8 and Scala 2.12.18 with Java 8 bytecode - replace Scala build plugin and match SLF4J 2 binding - supply embedded Spark module opens and bounded tests - isolate configurable server targets and verify Unicode - run connector CI on Java 11 and 17
- update the shared inventory from the published full reactor runtime - retain complete bundled license texts and required dependency notices - cover Jackson suffixed legal entries and the protobuf release license
- keep the Spark 3.5 SLF4J provider in provided scope - retain the exclusion of the inherited SLF4J 1 binding - document exact Guava acquisition and driver/executor paths - state the application classpath override and runtime boundary
- Inherit coverage tooling and keep scoped ORC access on Java 17. - Update test-only container support and include existing unit coverage. - Close both distinct clients and cover shared, separate and failing closes. - Preserve the current Java baseline; shared-base runtime validation is pending.
- remove the provided logging bridge from the runtime inventory - retain the exact published baseline Log4j API notice - verify the reactor copy and provider-free assembly legal coverage
- Apply the ORC module opening only on JDK 17 and later. - Keep default file-test selection when the JDK profile is active. - Preserve the primary close exception and suppress secondary failures. - Validate JVM option boundaries and unit/file suites on JDK 11 and 17.
- read plain backup streams without ZIP decoding - close invalid compressed streams before reporting errors - cover ZIP and plain backup restore round trips - allow a dedicated test server URL
- mark the validated submit example as client deploy mode - explain the remote driver Guava path in cluster deploy mode - distinguish driver provisioning from executor JAR copies
- Restore profile test selection without global file-test defaults. - Keep scoped ORC access and late coverage on JDK 11/17 test JVMs. - Mark client cleanup closed even when either client throws. - Cover primary, secondary and dual failures followed by retry.
- pass the configured URL to auth backup commands - pass the same URL to auth restore commands - keep command and client requests on one test fixture
Upgrade the backend to Spring Boot 3 and Jakarta APIs. Use a fresh H2 database with complete schema initialization. Align JAXB and Avatica runtime dependencies. Preserve runtime data while packaging archives and images. Carry the verified graph context and Java 17 documentation.
Run Hubble CI on Java 17 with Server 1.7 on Java 11. Use the shared pinned source checkout and compiler tooling. Keep server packaging separate from SDK dependency installation. Verify archive preservation before the distribution gate.
- update license selections for the Hubble Java 17 runtime - preserve required notices and complete bundled license texts - regenerate the shared published dependency inventory
Record the selected Server JDK before configuring Hubble Java 17. Use its explicit home for released Server build and startup. Keep Hubble package and acceptance on the module JVM.
- retain original module dependencies in the shared inventory - match the complete reactor dependency-copy result - verify existing license rows and published artifact hashes
- restore the complete notice for the retained published Jetty dependency - identify the exact artifact and original JAR entry in the main NOTICE - preserve every existing distribution notice unchanged
- distinguish the original gRPC runtime identity from Hubble additions - retain its existing Apache license selection and complete terms - align inventory classification with the full reactor classpaths
Bind Hikari settings before validating the effective connection. Check existing H2 metadata read-only before any schema initialization. Record the schema version only after fresh initialization succeeds. Verify old database bytes and rows remain unchanged on rejected startup.
- retain three complete bundled Jackson Core license resources - link the runtime inventory to the supplemental license texts - verify prefixed and suffixed legal entries against artifact hashes
Accept embedded file URLs with the optional file prefix omitted. Retain the native cipher when probing encrypted metadata read-only. Reject INIT and escaped setting names before opening a connection. Verify shorthand and AES restart plus unchanged encrypted legacy files.
Build the dependency license inventory on Java 17. Run CodeQL Java reactor compilation on Java 17. Preserve module target versions and fixture JVM settings.
- include the current tools and loader branches - preserve the independent shared and client stack - keep module runtime and source scopes unchanged
- include the current spark connector branch - retain the verified combined license inventory - provide the true multi-module cutover baseline # Conflicts: # hugegraph-dist/release-docs/LICENSE # hugegraph-dist/release-docs/NOTICE # hugegraph-dist/scripts/dependency/known-dependencies.txt
- prepare a Java 11 environment for the published server fixture - isolate server JAVA_HOME and PATH from the test JVM matrix - keep Java 11 and 17 CI results independent with fail-fast disabled
- use the merged ASF request encoding fix - align workflow and image source identities - retain full machine locks and short documentation
- pin the ASF binary release URL and official SHA-512 - verify release identity and both archive checksums on cache reuse - preserve Java 17 candidate builds and existing service startup - update offline fixture contracts for release downloads and failures
- fix the official 1.7.0 source repository and commit - reject mismatched callers before cache reads or downloads - validate and record release source identity in manifests - cover false source claims in offline fixture contracts
- reuse third-party Maven inputs without HugeGraph SDKs - remove the duplicate Hubble Maven cache - correct candidate build and Docker context instructions
keep Server configuration and Struct dependencies leave segmentation execution in the Server process avoid bundling the unused GPL Word dependency
imbajin
commented
Oct 7, 2026
imbajin
left a comment
Member
Author
There was a problem hiding this comment.
Blocking: no. Summary: The Spark API path still mishandles non-UTF-8 TEXT mappings and accepts incremental/recovery flags without implementing those modes. Evidence: exact-head reader and partition traces; Spark 3.5.8 documents UTF-8-only text input.
- download Server 1.7 from its current ASF location - retain the official checksum and source validation - update the fixture identity expectation
- find the unique runnable Server in official packages - preserve the two-instance configuration and tests - remove macOS sidecars before graph discovery
- reject non-UTF-8 TEXT mappings before Spark starts - reject unsupported incremental and failure modes - cover driver preflight in existing unit tests - document supported input and resume boundaries
- fall back only to the verified official archive - keep canonical cache identity and checksum failures - remove the unreachable short-ref rewrite
- detect tarfile extraction filter support before runtime setup - keep safe extraction mandatory without an unfiltered fallback - cover unsupported Python with existing archive unit tests
- remove restored Maven negative-cache markers - retain downloaded third-party dependencies - recheck repositories after transient failures
imbajin
commented
Oct 8, 2026
imbajin
left a comment
Member
Author
There was a problem hiding this comment.
Blocking: yes. Summary: The Spark API path can silently misdecode configured non-UTF-8 CSV sources, and the upstream compatibility workflow misses direct inputs in its PR trigger. Evidence: The inline comments identify the reader encoding and path-filter gaps at this head.
zyxxoo
previously approved these changes
Oct 8, 2026
Pengzna
previously approved these changes
Oct 8, 2026
- inherit the merged SDK build contract - inherit same-source image validation - keep Spark changes as the remaining PR diff
- reject non-UTF-8 CSV and JSON before Spark starts - cover FILE and HDFS preflight through real load calls - document UTF-8 input and preserve accepted aliases - include direct upstream compatibility trigger inputs
- propagate shared argument parser failures immediately - avoid command construction after missing option values - verify failures in the existing launcher test suite
This was referenced Oct 8, 2026
zyxxoo
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose of the PR
Spark Loader previously captured non-serializable partition state and mixed Loader and Spark submit options. It now uses partition-owned writers and explicit Java 17 application/runtime classpaths.
The core prerequisite #783 is merged. This branch is synchronized with the actual ASF master.
Module validation and product images use a fixed commit from ASF
apache/hugegraphmaster with Java 17 and TinkerPop 3.8.1. Machine source locks retain complete commit identities. Documentation and human-readable logs use short displays; standalone packaging validates the fixed identity locally without resolving a prefix through GitHub. An independent daily or manually triggered workflow checks the latest ASF master, resolving it once per run for SDK compilation and core HTTP compatibility tests.Main Changes
Validation
The pinned official Spark image is prepared in one environment step. The SDK baseline now includes the merged ASF UTF-8 request encoding fix from apache/hugegraph#3281. The US-ASCII Unicode readback check remains enabled. Updated-head CI is pending.
Coverage uploads remain in the existing test jobs without additional runners or artifact transfers. Upload errors are advisory, and project/patch coverage thresholds are informational in the root Codecov configuration.
The branch includes the latest ASF Toolchain master and preserves its independent license, advisory module validation, CodeQL and dependency-audit workflows. The SDK bootstrap retains flatten 1.3.0 and the explicit SDK dependency closure; it handles the same CI-friendly parent-version issue as the Server lifecycle fix while the reproducible baseline remains locked.
CSV/TEXT mappings require an explicit header and headerless input; the driver rejects unsupported header configurations before Spark initializes or submits partitions. Cluster mapping filenames must be unique among distributed resources; Spark retains its native collision checks.
The branch includes the merged #787 SDK contract and #784 image verification. Legacy CDC launcher argument failures now stop before command construction. The rebased #786 remains compatible with this branch, with no merge conflicts in the current simulated result.
The reproducible ASF Server baseline has advanced to include the merged UTF-8 encoding repair and the Maven parent-model fix. Machine identities remain complete, and documentation/log displays remain short. Runtime results for the updated baseline are pending CI.
CSV, JSON and TEXT mappings for FILE/HDFS require UTF-8 (including aliases). The driver rejects other charsets before Spark initialization or partition submission. The current-master compatibility trigger now also includes Java setup, Maven settings and Server startup inputs. Java 17 preflight and partition tests, launcher regressions and workflow lint passed locally; updated-head CI is pending.