Skip to content

feat: add support for dynamic KRaft quorum scaling - #1010

Draft
razvan wants to merge 9 commits into
mainfrom
feat/kraft-dynamic-voter-membership
Draft

feat: add support for dynamic KRaft quorum scaling#1010
razvan wants to merge 9 commits into
mainfrom
feat/kraft-dynamic-voter-membership

Conversation

@razvan

@razvan razvan commented Aug 18, 2026

Copy link
Copy Markdown
Member

Description

Fixes #1009

See CHANGELOG for details.

⌛ OKD integration tests

Outstanding issues

Issue 1: Controller bootstrap servers

Previously the property controller.quorum.bootstrap.servers used to contain a list of all controller pods. This meant that adding/removing a new controller would update this property in the configmap and the restart controller would restart ALL controller pods. This causes a lot of disruption and leads to long teardown times when deleting a CR or a kuttl namespace.

To avoid this, we now use a list of headless services (one per role group) which are much more stable as there is usually only one of them. But this uncovered a KRaft bug where a new controller tries to fetch KRaft metadata from its self and crashloops with:

kafka 2026-08-19T12:42:23,080 INFO [kafka-2110489704-metadata-loader-event-handler] org │
│ .apache.kafka.image.loader.MetadataLoader - [MetadataLoader id=2110489704] initializeNe │
│ wPublishers: the loader is still catching up because we still don't know the high water │
│ mark yet.

Workarount: implemented a livenessProbe that fails if the controller stays in state unattached for too long.

Issue 2: CR deletion

When a KAfka CR is deleted, the controllers may shut down faster than the brokers leaving them hanging.

This is especially a problem for kuttl tests.

Workaround: a preStop sleep.

Definition of Done Checklist

  • Not all of these items are applicable to all PRs, the author should update this template to only leave the boxes in that are relevant
  • Please make sure all these things are done and tick the boxes

Author

  • Changes are OpenShift compatible
  • CRD changes approved
  • CRD documentation for all fields, following the style guide.
  • Helm chart can be installed and deployed operator works
  • Integration tests passed (for non trivial changes)
  • Changes need to be "offline" compatible
  • Links to generated (nightly) docs added
  • Release note snippet added

Reviewer

  • Code contains useful comments
  • Code contains useful logging statements
  • (Integration-)Test cases added
  • Documentation added or updated. Follows the style guide.
  • Changelog updated
  • Cargo.toml only contains references to git tags (not specific commits or branches)

Acceptance

  • Feature Tracker has been updated
  • Proper release label has been added
  • Links to generated (nightly) docs added
  • Release note snippet added
  • Add type/deprecation label & add to the deprecation schedule
  • Add type/experimental label & add to the experimental features tracker

razvan and others added 6 commits August 18, 2026 14:24
Controllers now run a quorum-manager sidecar that admits itself into the
KRaft voter set on startup (add-controller) and removes itself before
termination (remove-controller via preStop), so controller role groups
can be scaled up/down on a running cluster without a full rolling
restart. controller.quorum.bootstrap.servers now points at each
controller role group's headless Service DNS name instead of individual
pod addresses, keeping container commands stable across replica changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Controllers get a startupProbe (plain TCP, generous failure threshold
for slow metadata-log replay on boot), a plain-TCP livenessProbe, and a
readinessProbe that checks the node's Raft state via its metrics
endpoint instead of a bare TCP check, so a controller stuck rejoining
the quorum is correctly reported as not ready.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Controller pods now start/scale sequentially (OrderedReady) instead of
in parallel, since the quorum-manager sidecar's admission flow assumes
one voter joins at a time. Brokers are unaffected and keep Parallel.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scaling controllers to 0 replicas while brokers keep running is now
rejected at validation time with an actionable error, instead of
failing much later and confusingly while building the broker's
ConfigMap. Scaling controllers and brokers to 0 together (a coordinated
whole-cluster stop) is still allowed and now actually builds, since
downstream resource builders no longer assume a non-empty controller
quorum whenever KRaft mode is active.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ns tests

Enable the previously version-gated scale-up/down steps (Kafka 3.7 no
longer needs special-casing), assert quorum voter counts via
kafka-metadata-quorum.sh after each scale, and add a final step scaling
both controllers and brokers to 0 to exercise the whole-cluster-stop
path before namespace teardown.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Update the KRaft controller usage guide for scale-up/down support,
record the design spec and implementation plan, add the CHANGELOG
entries for this branch's changes, and ignore .worktrees/ for local
worktree checkouts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@razvan razvan self-assigned this Aug 18, 2026
@razvan

razvan commented Aug 18, 2026

Copy link
Copy Markdown
Member Author

The failed test was due to namespace deletion timeout

--- FAIL: kuttl (2409.35s)
    --- FAIL: kuttl/harness (0.00s)
        --- PASS: kuttl/harness/logging_kafka-3.9.2_zookeeper-latest-3.9.5_openshift-false (134.90s)
        --- FAIL: kuttl/harness/smoke-kraft_kafka-kraft-4.2.1_openshift-false (634.74s)
        --- PASS: kuttl/harness/smoke_kafka-3.9.2_zookeeper-3.9.5_use-client-tls-false_openshift-false (81.56s)
        --- PASS: kuttl/harness/upgrade_upgrade_old-3.9.2_upgrade_new-4.2.1_use-client-tls-false_use-client-auth-tls-false_openshift-false (120.71s)
        --- PASS: kuttl/harness/kerberos_kafka-3.9.2_zookeeper-latest-3.9.5_openshift-false_krb5-1.21.1_kerberos-realm-PROD.MYCORP_kerberos-backend-mit_broker-listener-class-cluster-internal_bootstrap-listener-class-external-unstable (113.64s)
        --- PASS: kuttl/harness/tls_kafka-3.9.2_zookeeper-latest-3.9.5_use-client-tls-false_use-client-auth-tls-false_openshift-false (33.64s)
        --- PASS: kuttl/harness/cluster-operation_kafka-latest-3.9.2_zookeeper-latest-3.9.5_openshift-false (41.49s)
        --- PASS: kuttl/harness/configuration_kafka-latest-3.9.2_openshift-false (11.45s)
        --- PASS: kuttl/harness/operations-kraft_kafka-kraft-4.2.1_openshift-false (1079.48s)
        --- PASS: kuttl/harness/opa_kafka-latest-3.9.2_zookeeper-latest-3.9.5_opa-latest-1.16.2_use-opa-tls-false_openshift-false_krb5-1.21.1 (128.33s)
        --- PASS: kuttl/harness/delete-rolegroup_kafka-3.9.2_zookeeper-latest-3.9.5_openshift-false (29.40s)
FAIL
ERROR:root:kuttl failed

@razvan
razvan marked this pull request as draft August 18, 2026 15:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feat: KRaft quorum scalability

1 participant