Skip to content

cuda.bindings: support multiple CTK release lines on main - #2737

Open
rwgk wants to merge 31 commits into
NVIDIA:mainfrom
rwgk:agent/cuda-bindings-12-on-main
Open

cuda.bindings: support multiple CTK release lines on main#2737
rwgk wants to merge 31 commits into
NVIDIA:mainfrom
rwgk:agent/cuda-bindings-12-on-main

Conversation

@rwgk

@rwgk rwgk commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

REMINDER

Before merging, remove the temporary .lycheeignore before triggering final CI. It excludes only three canonical main/cuda_bindings_12 URLs that cannot resolve until this PR is merged. The authored-source lychee hook is skipped by CI, so removing the file will not prevent final CI from passing.

After merging, run pre-commit run lychee --all-files on fresh main to validate those links.

Summary

Closes #1199.

This PR is the writable continuation of Keith Kraus's original PR #2675, "cuda.bindings: build 12.9 and 13.x selectively from main". Most of the CUDA 12 source import and the initial build, test, and release integration originated in Keith's PR.

PR #2675 was automatically closed when its temporary pull-request/2467 base was deleted after #2467 merged. Its head branch was not maintainer-writable, so this replacement preserves that work and commit history, retargets it to current main, and completes the redesign requested during review.

The result is one active development branch for both released CUDA bindings lines:

  • CUDA 12.9 and CUDA 13.3 bindings are built, tested, documented, and released from main.
  • ci/versions.yml is the authoritative release-line registry for CI and release tooling.
  • The historical 12.9.x branch becomes a read-only release record, not an active backport or artifact-source branch.
  • The obsolete backport workflow is removed.

This builds on the dependency-aware selective CI merged in #2467.

Release-Line Registry

ci/versions.yml records the two released bindings lines:

Line ID Role CTK target Source root Release-tag family
released-12 maintenance 12.9 cuda_bindings_12/ v12.9.*
released-13 current 13.3 cuda_bindings/ v13.3.*

Each line also declares its exact toolkit pin and channel and its prerelease-tag policy.

ci/tools/bindings_config.py validates the registry and emits normalized records for downstream consumers. The registry has one current line and an ordered maintenance list. Every configured line must be assigned exactly one role.

CI and release logic select stable line IDs or roles instead of treating CUDA versions, source-directory names, and current/backport as interchangeable concepts. Updating a role is therefore centralized and reviewable.

The monolithic wheel builder retains one explicit transitional boundary: it currently requires exactly one current line and one maintenance line with different CUDA ABI majors. Unsupported registry shapes fail closed instead of producing incomplete artifacts.

Source Layout and Maintenance Model

The filesystem names identify their contents; the registry identifies their orchestration role:

  • cuda_bindings/ contains the released CUDA 13 line.
  • cuda_bindings_12/ contains the released CUDA 12.9 line.
  • current and maintenance exist only as registry roles, not as directory aliases.

The two complete package roots are an intentional transitional design. Most of cuda_bindings_12/ is imported from NVIDIA/cuda-python@238955935bd903ac72817c0dfdfe4f6a54ee6bb1:cuda_bindings. cuda_bindings_12/MAINTENANCE.md records ownership and generation provenance, including the portion reproduced by cybind commit 95d8bb525de46a9ff7ae40d759a98cbe50cf8391.

ci/cuda-bindings-shared-files.json lists the small handwritten subset that must remain byte-identical across the two roots, and a pre-commit/CI checker enforces that invariant. Generated and cybind-owned support files are guarded by their existing content seals and recorded provenance rather than being required to match across different CTK targets.

For subsequent bindings fixes, contributors must update every applicable root or document concretely why one line is unaffected.

Build, Test, and Release Behavior

Change or event Result
Source change under one registered bindings root Build and test that bindings line and its matching cuda-python package; exercise the relevant CUDA Core ABI as needed
Shared bindings consumer or CI/release infrastructure change Exercise both released lines affected by the shared change
v12.9.* release tag Select the registered CUDA 12.9 bindings/metapackage pair
v13.3.* release tag Select the registered CUDA 13.3 bindings/metapackage pair

Release selection uses the registry stored in the tagged source tree. A compatibility path supports older tags whose source trees predate the registry and still use the generic cuda_bindings/ directory.

One commit may carry one CUDA 12.9 release tag and one CUDA 13.3 release tag. Each tag independently selects the matching line, package root, metadata, and artifacts. Each tag still triggers its own release run; combining both releases into one run is not required here. Two different releases from the same tag family should not be placed on one commit.

CUDA 12 development versions advance normally after the latest stable CUDA 12 tag becomes reachable instead of remaining pinned indefinitely to the same development version.

Decisions Requested From Reviewers

Please explicitly accept or reject these policies:

  1. main is the sole active source of truth. The historical 12.9.x branch receives no further routine or emergency backports. Applicable CUDA 12 fixes are made in cuda_bindings_12/ on main, alongside any corresponding current-line change.
  2. Released bindings use explicit source roots. Directory names identify their contents; orchestration roles are centralized in ci/versions.yml.
  3. Full-root duplication is transitional. The identical-file checker is a short-term safety mechanism for the small shared handwritten subset, not the preferred long-term sharing model.

The generated NVML memoryview change discussed during the first review is not part of this PR and should be handled independently if needed.

Review Map

The high-value review surface is outside the imported CUDA 12 tree:

  • Registry and semantics: ci/versions.yml, ci/tools/bindings_config.py, and their tests
  • Selective planning: ci/tools/compute_ci_plan.py, matrix validation, and their tests
  • Build and test routing: GitHub workflows plus artifact, environment, and wheel helpers under ci/tools/
  • Release routing and versioning: release workflows, tag resolution, release-note validation, and maintenance-line SCM handling
  • Cross-root safety: ci/cuda-bindings-shared-files.json, its checker, and cuda_bindings_12/MAINTENANCE.md

Most files under cuda_bindings_12/ are the direct CUDA 12.9 import from Keith's PR #2675. Differences from cuda_bindings/ are generally target- or generation-specific and intentional.

Validation

  • TestVenv/bin/python -m pytest ci/tools/tests: 161 passed on final local head c6a0cf1
  • pre-commit run --all-files, including registry validation, shared-file checks, generated-file seals, Ruff, actionlint, YAML/TOML/RST checks, and lychee with the temporary exception described in the REMINDER
  • Local source-build, import, extension-build, dependency-metadata, and environment-routing validation for both registered CUDA toolkits
  • Release-tag and tagged-tree tests covering independent CUDA 12.9 and CUDA 13.3 selection when both tag families point to one commit
  • Full gated GitHub Actions matrix on review head c6a0cf1: run 33572804123 passed with all 104 jobs successful

Out of Scope

  • Refactoring generated target-specific content into overlays or otherwise eliminating the two complete package roots
  • Generalizing the monolithic wheel builder beyond the two-line shape supported here
  • Combining multiple release tags into one release workflow run
  • The separate generated NVML memoryview change discussed during review

Checklist

  • New or existing tests cover these changes.
  • Documentation is updated for the maintenance and release-line model.

@rwgk rwgk added this to the cuda.bindings 13.5.0 & 12.9.10 milestone Aug 31, 2026
@rwgk rwgk added enhancement Any code-related improvements CI/CD CI/CD infrastructure cuda.bindings Everything related to the cuda.bindings module labels Aug 31, 2026
@rwgk rwgk self-assigned this Aug 31, 2026
@copy-pr-bot

copy-pr-bot Bot commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

rwgk commented Aug 31, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test b87d0a1

@github-actions

Copy link
Copy Markdown

@mdboom mdboom left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I started to comment on some individual things, but then decided to stop because I think there is a more fundamental change that needs to be made across this whole PR (and then I'm happy to come back and review further).

/Today/ the "current" version is 13, and the backport version is 12. But at some point in the future that will switch to 14 and 13. This pervasively hardcodes those version numbers all over this codebase, especially in CI, but in a bunch of the release scripts as well, and even the cuda_bindings_12 directory name as indicators of current vs. backport.

Instead, we should use the config we already have in versions.yml and use that to drive the numbers everywhere. That way when it's time to move on, all that should be required is updating versions.yml, and copying/overwriting the existing cuda_bindings to cuda_bindings_backport (or whatever we want to call it), and move on. I'm sure there are many details I'm missing, but that should be the goal and design -- it would be preferable to reduce it to as close to that as possible. The problem with this as-is is that there are hundreds of context-sensitive places that would need to be updated to do that update -- we are creating a massive pile of technical debt to pay later. I'm sure an agent might get that X% correct, but I always think it's better to engineer for flexibility, especially for something we know will happen. If versions.yml (which requires using yq to parse etc.) makes this too difficult, we could explore a simple VARIABLE=value format which would parse as both bash variables and Python variables and probably be more convenient to use from the many places it is needed. There are really only two actual values in versions.yml today, so that should be fine.

I'm also a little concerned (without any testing-based evidence) that this will break when we tag the same commit with v13.x.y and v12.x.y, which will be the common case, in fact, IMHO, one of the real benefits of moving to this approach. We should get an agent to do a thorough investigation of that use case and make sure it is covered. Ideally, it would be nice for a single release run to do both releases simultaneously but it's not a deal breaker if it still requires kicking off two runs.

Also what is this (from the agent's PR description):

The later NVML memoryview fix is reproduced byte-for-byte from cybind commit
6def52ca508c9e14ef67f4ce26a0c677f3fbad72 with Doxygen 1.17.0:

If there is something like this that wasn't backported, let's deal with that separately so it's not an unrelated tag-along to this PR.

Also a note for future agent reviewers of this PR: The interesting part of this PR is the part outside of the cuda_bindings_backport or cuda_bindings_12 directory. Those are just direct copies from the 12.9.x branch, and any differences between that and the cuda_bindings directory are likely intentional. When reviewing, focus on the scaffolding / CI / overall structure.

Comment thread .github/workflows/backport.yml Outdated
Comment thread .github/workflows/build-wheel.yml Outdated
Comment thread ci/cuda-bindings-shared-files.json Outdated
Comment thread ci/cuda-bindings-shared-files.json Outdated
Comment thread ci/ci-pipeline.svg
Comment thread cuda_python/docs/environment-docs.yml
Comment thread ci/tools/check_cuda_bindings_shared_files.py
@rwgk rwgk changed the title cuda.bindings: build 12.9 and 13.x selectively from main cuda.bindings: support multiple CTK release lines on main Aug 31, 2026
@rwgk

rwgk commented Aug 31, 2026

Copy link
Copy Markdown
Contributor Author

Archiving options related to a lychee chicken-and-egg issue. I'll go with Option 1 below. This comment is to explain why.


codex:

We have three sensible options. For PR 2737, I recommend keeping the canonical links unchanged and treating these as documented pre-merge exceptions.

  1. One-off exact exclusions — recommended

Run lychee once with only these three URLs excluded, record that every other link passes, and rerun without exclusions after merge.

This is reasonable because authored-source lychee is explicitly skipped by the GitHub CI job at .github/workflows/ci.yml, so these are not merge-gating failures. It avoids landing temporary configuration or compromising the final URLs.

  1. Temporary .lycheeignore

Add three exact anchored patterns so pre-commit run --all-files is completely green, then remove them immediately after merging. Lychee officially supports checked-in, commented exclusions via .lycheeignore. Lychee exclusion documentation

This is practical, but creates a mandatory cleanup PR and briefly leaves three blind spots on main.

  1. Permanent local remapping

Teach the hook to map:

https://github.com/NVIDIA/cuda-python/{tree,blob}/main/<path>
→ file://<current-worktree>/<path>

Then links to newly introduced files are validated against the checkout before they exist online. This is exactly the future-URL use case for lychee’s remapping feature. Lychee remapping documentation

It is the principled reusable solution, but needs a portable wrapper to calculate the absolute worktree path. I would pursue it separately only if this problem starts recurring.

I would avoid:

  • Accepting all 404 responses; lychee’s setting is global.
  • Excluding the entire cuda_bindings_12 subtree or source files.
  • Replacing the links with fork, PR, SHA, or repository-root URLs; those are inferior permanent package metadata.
  • Splitting out a bootstrap PR solely to make the paths exist.

So my recommendation is option 1: preserve the three correct final URLs, validate everything else with exact one-off exclusions, and rerun lychee from fresh main after merging.

@github-actions github-actions Bot added cuda.core Everything related to the cuda.core module cuda.pathfinder Everything related to the cuda.pathfinder module labels Sep 1, 2026
@rwgk

rwgk commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test f4ddc4e

@rwgk

rwgk commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test c6a0cf1

@rwgk
rwgk marked this pull request as ready for review September 2, 2026 03:35
@rwgk

rwgk commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test 83c1cf0

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD CI/CD infrastructure cuda.bindings Everything related to the cuda.bindings module cuda.core Everything related to the cuda.core module cuda.pathfinder Everything related to the cuda.pathfinder module enhancement Any code-related improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Revisit cuda-bindings branching strategy

3 participants