Skip to content

refactor(store): drain TTL native resources on shutdown - #3304

Merged
imbajin merged 20 commits into
apache:masterfrom
hugegraph:fix/store-shutdown-native-drain
Oct 10, 2026
Merged

imbajin merged 20 commits into
apache:masterfrom
hugegraph:fix/store-shutdown-native-drain

Conversation

@imbajin

@imbajin imbajin commented Oct 9, 2026 •

Copy link
Copy Markdown
Member
image Shutdown waits for TTL workers and their native cleanup after RPC/Scan ownership drains. TTL closes its scan and borrowed session lease in order, attempts both releases, and retains cleanup failures as shutdown blockers.

The distribution stop script reports timeouts explicitly. Native reference-count, close-failure, active-read and shutdown regressions ship here. A real nonempty TTL scan with interrupted Spring close verifies that the database stays alive while the reader is blocked, then checks iterator/session drain before bean destruction.

This branch inherits the RPC final-batch, one-shot cancellation and aggregate UNAVAILABLE fixes. Callback-failure shutdown fixtures follow the error terminal while retaining cleanup. Independent review and the complete local lifecycle/shutdown suite passed; backend integration results are reported by the PR checks. Full process exit/recovery and scan benchmarks remain separate validation boundaries.

TTL and stop behavior extend docs/storage-lifecycle.md, the same guide as RPC/Scan. README links to the canonical section. Depends on RPC/Scan; Server remains independent. No second ownership registry is introduced.

- align client cancellation and server half-close behavior
- publish final batches outside response parsing locks
- drain RPC callbacks and scan owners before storage close
- retain TTL cleanup failures and finish interrupted Raft joins
- preserve executable packaging and lifecycle CI coverage
- retain parser failures across completion races
- drain heartbeat producers despite interruption
- restore interrupts after native database teardown
- cover real native close in an isolated subprocess
- keep client and server request lifecycle changes together
- retain scoped scan and query cleanup before database close
- partition lifecycle tests without changing their assertions
- defer global RPC and TTL shutdown to a dependent PR
- close global RPC admission before request cleanup waits
- retain batch half-close credit and explicit client cancellation
- track unary query and parallel count resource owners
- preserve failed native cleanup independently of RPC completion
- keep matching regressions in the RPC prerequisite
@codecov

codecov Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 41.73%. Comparing base (72085aa) to head (a40f64d).
⚠️ Report is 1 commits behind head on master.

Additional details and impacted files
@@             Coverage Diff              @@
##             master    #3304      +/-   ##
============================================
+ Coverage     41.71%   41.73%   +0.02%     
- Complexity     7033     7042       +9     
============================================
  Files           762      762              
  Lines         67047    67047              
  Branches       9018     9018              
============================================
+ Hits          27966    27981      +15     
+ Misses        35872    35857      -15     
  Partials       3209     3209              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

- retain off-thread dispatch during executor saturation
- verify cancellation and drain through real Netty RPCs
- document callback queue and shutdown boundaries
- interrupt active readers before iterator cleanup waits
- verify context cancellation and deadline cleanup
- propagate half-close worker failures into assertions
@imbajin
imbajin force-pushed the fix/store-shutdown-native-drain branch from c221fe1 to 0a3f050 Compare October 9, 2026 21:00
@imbajin imbajin changed the title fix(store): drain global RPC and TTL resources fix(store): drain TTL native resources on shutdown Oct 9, 2026
- listen to context cancellation during blocked unary work
- remove listeners while retaining iterator cleanup accounting
- verify query and count cancellation through real Netty RPCs
@imbajin
imbajin force-pushed the fix/store-shutdown-native-drain branch from a31092a to 813448a Compare October 9, 2026 21:50
- interrupt owned batch readers before iterator locks
- keep closed partition iterators terminal
- refresh the timeout when accepting a response
- cover blocked shutdown and parsing boundaries
@imbajin
imbajin force-pushed the fix/store-shutdown-native-drain branch from 813448a to 663b8bc Compare October 9, 2026 22:08
- combine RPC, recovery and provider guidance
- preserve WAL and native ownership safety contracts
- remove PR-stage notes and repeated explanations
- update entry points to the canonical guide
- recheck the queue after an empty timed poll
- preserve producer publication before completion
- cover final-batch delivery at the poll boundary
- extend the RPC prerequisite with sticky TTL cleanup ownership
- wait for TTL native release before database destruction
- preserve stop timeout diagnostics and PID evidence
- keep application shutdown tests with their incremental behavior
- release scans before borrowed native sessions
- retain failures while attempting every lease cleanup
- verify native refcounts and fix shutdown test imports
- retain raw native handles for post-close assertions
- model native reads that outlast cancellation interrupts
- verify iterator close waits for actual read completion
- use the existing Ubuntu host for mocked stop checks
- avoid unnecessary image pulls and rate-limit failures
- retain timeout and PID-retention assertions
- keep README as a concise shutdown entry point
- preserve operational details in store-shutdown.md
- remove duplicate timeout and drain instructions
- keep stop instructions in the storage main guide
- replace the removed shutdown document reference
- retain a concise README entry point
- exercise a nonempty native TTL scan during close
- keep the database alive while cleanup is blocked
- preserve interrupted Spring shutdown semantics
- verify iterator leases drain before bean destruction
@imbajin
imbajin force-pushed the fix/store-shutdown-native-drain branch from 663b8bc to cb7e120 Compare October 10, 2026 04:34
Interrupt active one-shot callers and preserve thread reuse.
Close expired streams outside the admission monitor.
Report aggregate shutdown cancellation as unavailable.
Cover cancellation, terminal delivery and retained cleanup.
Keep shutdown cancellation consistent with the RPC prerequisite.
Retain active native ownership until cleanup completes.
Adapt callback-failure regressions to unavailable terminals.
- merge the published RPC and server lifecycle changes
- retain TTL scheduler, worker and native cleanup drains
- preserve shutdown tests, CI and lifecycle guidance
@imbajin imbajin changed the title fix(store): drain TTL native resources on shutdown refactor(store): drain TTL native resources on shutdown Oct 10, 2026
@imbajin
imbajin merged commit ed3a2df into apache:master Oct 10, 2026
35 checks passed
@imbajin
imbajin deleted the fix/store-shutdown-native-drain branch October 10, 2026 10:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants