Skip to content

feat(loader): upgrade Java 17 for Spark-API - #785

Merged
imbajin merged 244 commits into
apache:masterfrom
hugegraph:java17-spark
Oct 8, 2026
Merged

imbajin merged 244 commits into
apache:masterfrom
hugegraph:java17-spark

Conversation

@imbajin

@imbajin imbajin commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Purpose of the PR

Spark Loader previously captured non-serializable partition state and mixed Loader and Spark submit options. It now uses partition-owned writers and explicit Java 17 application/runtime classpaths.

Spark Java 17 input preflight and partition-owned HTTP writers

The core prerequisite #783 is merged. This branch is synchronized with the actual ASF master.

Module validation and product images use a fixed commit from ASF apache/hugegraph master with Java 17 and TinkerPop 3.8.1. Machine source locks retain complete commit identities. Documentation and human-readable logs use short displays; standalone packaging validates the fixed identity locally without resolving a prefix through GitHub. An independent daily or manually triggered workflow checks the latest ASF master, resolving it once per run for SDK compilation and core HTTP compatibility tests.

Main Changes

  • Serialize only partition inputs and restore the supported structured CSV/TEXT parsing path.
  • Preserve dry-run without graph mutations and propagate invalid option failures as unsuccessful process exits.
  • Forward extra drivers through explicit --jars. Default to client deployment; cluster requires a command-line deploy mode and master. Reject standalone cluster deployment and cluster mapping paths containing commas or fragments.
  • Run the packaged API loading job with a separate executor and verify IDs, Unicode values and failure status.

Validation

The pinned official Spark image is prepared in one environment step. The SDK baseline now includes the merged ASF UTF-8 request encoding fix from apache/hugegraph#3281. The US-ASCII Unicode readback check remains enabled. Updated-head CI is pending.

Coverage uploads remain in the existing test jobs without additional runners or artifact transfers. Upload errors are advisory, and project/patch coverage thresholds are informational in the root Codecov configuration.

The branch includes the latest ASF Toolchain master and preserves its independent license, advisory module validation, CodeQL and dependency-audit workflows. The SDK bootstrap retains flatten 1.3.0 and the explicit SDK dependency closure; it handles the same CI-friendly parent-version issue as the Server lifecycle fix while the reproducible baseline remains locked.

CSV/TEXT mappings require an explicit header and headerless input; the driver rejects unsupported header configurations before Spark initializes or submits partitions. Cluster mapping filenames must be unique among distributed resources; Spark retains its native collision checks.

The branch includes the merged #787 SDK contract and #784 image verification. Legacy CDC launcher argument failures now stop before command construction. The rebased #786 remains compatible with this branch, with no merge conflicts in the current simulated result.

The reproducible ASF Server baseline has advanced to include the merged UTF-8 encoding repair and the Maven parent-model fix. Machine identities remain complete, and documentation/log displays remain short. Runtime results for the updated baseline are pending CI.

CSV, JSON and TEXT mappings for FILE/HDFS require UTF-8 (including aliases). The driver rejects other charsets before Spark initialization or partition submission. The current-master compatibility trigger now also includes Java setup, Maven settings and Server startup inputs. Java 17 preflight and partition tests, launcher regressions and workflow lint passed locally; updated-head CI is pending.

imbajin added 30 commits October 3, 2026 23:53
- manage compiler, surefire and coverage plugin versions
- keep Java 8 output and use compatible Mockito 4
- resolve optional server repositories, refs and pins
- verify checkout and package servers without install
- share authenticated startup and filter AppleDouble files
- connect repository and fetch refs to all server installers
- reuse the shared installer for Tools and Spark validation
- preserve Loader source caching with verified checkouts
- isolate fixture builds and retain separate HTTPS RPC ports
- preserve Mockito 2 interaction assertions with the equivalent API
- align test-only Byte Buddy dependencies with Mockito 4
- retain the Spring Boot 2 and Java 11 shared baseline
- run shared server selection and installer tests in one CI job
- document Maven 3.6.3 as the minimum source build version
- report missing or ambiguous server archives before exiting
- trigger push checks for client installer changes
- trigger pull request checks for the same source path
- cover the installer exercised by the existing contracts
- encode legacy graph configuration as properties
- keep task and target tests aligned with server contracts
- inherit shared compiler and coverage plugin versions
- preserve Java 8 APIs in compatibility tests
- align Spark 3.5.8 and Scala 2.12.18 with Java 8 bytecode
- replace Scala build plugin and match SLF4J 2 binding
- supply embedded Spark module opens and bounded tests
- isolate configurable server targets and verify Unicode
- run connector CI on Java 11 and 17
- update the shared inventory from the published full reactor runtime
- retain complete bundled license texts and required dependency notices
- cover Jackson suffixed legal entries and the protobuf release license
- keep the Spark 3.5 SLF4J provider in provided scope
- retain the exclusion of the inherited SLF4J 1 binding
- document exact Guava acquisition and driver/executor paths
- state the application classpath override and runtime boundary
- Inherit coverage tooling and keep scoped ORC access on Java 17.
- Update test-only container support and include existing unit coverage.
- Close both distinct clients and cover shared, separate and failing closes.
- Preserve the current Java baseline; shared-base runtime validation is pending.
- remove the provided logging bridge from the runtime inventory
- retain the exact published baseline Log4j API notice
- verify the reactor copy and provider-free assembly legal coverage
- Apply the ORC module opening only on JDK 17 and later.
- Keep default file-test selection when the JDK profile is active.
- Preserve the primary close exception and suppress secondary failures.
- Validate JVM option boundaries and unit/file suites on JDK 11 and 17.
- read plain backup streams without ZIP decoding
- close invalid compressed streams before reporting errors
- cover ZIP and plain backup restore round trips
- allow a dedicated test server URL
- mark the validated submit example as client deploy mode
- explain the remote driver Guava path in cluster deploy mode
- distinguish driver provisioning from executor JAR copies
- Restore profile test selection without global file-test defaults.
- Keep scoped ORC access and late coverage on JDK 11/17 test JVMs.
- Mark client cleanup closed even when either client throws.
- Cover primary, secondary and dual failures followed by retry.
- pass the configured URL to auth backup commands
- pass the same URL to auth restore commands
- keep command and client requests on one test fixture
Upgrade the backend to Spring Boot 3 and Jakarta APIs.
Use a fresh H2 database with complete schema initialization.
Align JAXB and Avatica runtime dependencies.
Preserve runtime data while packaging archives and images.
Carry the verified graph context and Java 17 documentation.
Run Hubble CI on Java 17 with Server 1.7 on Java 11.
Use the shared pinned source checkout and compiler tooling.
Keep server packaging separate from SDK dependency installation.
Verify archive preservation before the distribution gate.
- update license selections for the Hubble Java 17 runtime
- preserve required notices and complete bundled license texts
- regenerate the shared published dependency inventory
Record the selected Server JDK before configuring Hubble Java 17.
Use its explicit home for released Server build and startup.
Keep Hubble package and acceptance on the module JVM.
- retain original module dependencies in the shared inventory
- match the complete reactor dependency-copy result
- verify existing license rows and published artifact hashes
- restore the complete notice for the retained published Jetty dependency
- identify the exact artifact and original JAR entry in the main NOTICE
- preserve every existing distribution notice unchanged
- distinguish the original gRPC runtime identity from Hubble additions
- retain its existing Apache license selection and complete terms
- align inventory classification with the full reactor classpaths
Bind Hikari settings before validating the effective connection.
Check existing H2 metadata read-only before any schema initialization.
Record the schema version only after fresh initialization succeeds.
Verify old database bytes and rows remain unchanged on rejected startup.
- retain three complete bundled Jackson Core license resources
- link the runtime inventory to the supplemental license texts
- verify prefixed and suffixed legal entries against artifact hashes
Accept embedded file URLs with the optional file prefix omitted.
Retain the native cipher when probing encrypted metadata read-only.
Reject INIT and escaped setting names before opening a connection.
Verify shorthand and AES restart plus unchanged encrypted legacy files.
Build the dependency license inventory on Java 17.
Run CodeQL Java reactor compilation on Java 17.
Preserve module target versions and fixture JVM settings.
- include the current tools and loader branches
- preserve the independent shared and client stack
- keep module runtime and source scopes unchanged
- include the current spark connector branch
- retain the verified combined license inventory
- provide the true multi-module cutover baseline

# Conflicts:
#	hugegraph-dist/release-docs/LICENSE
#	hugegraph-dist/release-docs/NOTICE
#	hugegraph-dist/scripts/dependency/known-dependencies.txt
- prepare a Java 11 environment for the published server fixture
- isolate server JAVA_HOME and PATH from the test JVM matrix
- keep Java 11 and 17 CI results independent with fail-fast disabled
- use the merged ASF request encoding fix
- align workflow and image source identities
- retain full machine locks and short documentation
- pin the ASF binary release URL and official SHA-512
- verify release identity and both archive checksums on cache reuse
- preserve Java 17 candidate builds and existing service startup
- update offline fixture contracts for release downloads and failures
- fix the official 1.7.0 source repository and commit
- reject mismatched callers before cache reads or downloads
- validate and record release source identity in manifests
- cover false source claims in offline fixture contracts
- reuse third-party Maven inputs without HugeGraph SDKs
- remove the duplicate Hubble Maven cache
- correct candidate build and Docker context instructions
keep Server configuration and Struct dependencies

leave segmentation execution in the Server process

avoid bundling the unused GPL Word dependency

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: no. Summary: The Spark API path still mishandles non-UTF-8 TEXT mappings and accepts incremental/recovery flags without implementing those modes. Evidence: exact-head reader and partition traces; Spark 3.5.8 documents UTF-8-only text input.

- download Server 1.7 from its current ASF location
- retain the official checksum and source validation
- update the fixture identity expectation
- find the unique runnable Server in official packages
- preserve the two-instance configuration and tests
- remove macOS sidecars before graph discovery
- reject non-UTF-8 TEXT mappings before Spark starts
- reject unsupported incremental and failure modes
- cover driver preflight in existing unit tests
- document supported input and resume boundaries
- fall back only to the verified official archive
- keep canonical cache identity and checksum failures
- remove the unreachable short-ref rewrite
- detect tarfile extraction filter support before runtime setup
- keep safe extraction mandatory without an unfiltered fallback
- cover unsupported Python with existing archive unit tests
- remove restored Maven negative-cache markers
- retain downloaded third-party dependencies
- recheck repositories after transient failures

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The Spark API path can silently misdecode configured non-UTF-8 CSV sources, and the upstream compatibility workflow misses direct inputs in its PR trigger. Evidence: The inline comments identify the reader encoding and path-filter gaps at this head.

Comment thread .github/workflows/upstream-compatibility.yml
zyxxoo
zyxxoo previously approved these changes Oct 8, 2026
Pengzna
Pengzna previously approved these changes Oct 8, 2026
- inherit the merged SDK build contract
- inherit same-source image validation
- keep Spark changes as the remaining PR diff
- reject non-UTF-8 CSV and JSON before Spark starts
- cover FILE and HDFS preflight through real load calls
- document UTF-8 input and preserve accepted aliases
- include direct upstream compatibility trigger inputs
@imbajin
imbajin dismissed stale reviews from Pengzna and zyxxoo via e383ed2 October 8, 2026 12:13
- propagate shared argument parser failures immediately
- avoid command construction after missing option values
- verify failures in the existing launcher test suite
@imbajin imbajin changed the title fix(loader): run Spark API loading on Java 17 feat(loader): upgrade Java 17 for Spark-API Oct 8, 2026
@imbajin
imbajin merged commit d49fc5f into apache:master Oct 8, 2026
31 of 32 checks passed
@imbajin
imbajin deleted the java17-spark branch October 8, 2026 16:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

loader hugegraph-loader

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants