Skip to content

Fix cuda.coop limitation preventing user-defined types when items_per_thread > 1 in block scan module. - #4756

Merged
tpn merged 4 commits into
NVIDIA:mainfrom
tpn:4747-fix-items_per_thread-restriction-in-cuda-coop
May 21, 2025
Merged

Fix cuda.coop limitation preventing user-defined types when items_per_thread > 1 in block scan module.#4756
tpn merged 4 commits into
NVIDIA:mainfrom
tpn:4747-fix-items_per_thread-restriction-in-cuda-coop

Conversation

@tpn

@tpn tpn commented May 20, 2025

Copy link
Copy Markdown
Contributor

Commits have been organized to ease review. Underlying "bug" was fixed by adding a missing lower_constant for the Complex type we use in our tests of user-defined types. _block_scan.py invariants and docstrings updated accordingly.

Additionally, now that there are no restrictions regarding items_per_thread > 1, I have changed the default from 1 to 4, such that we have a more optimal default out-of-the-box.

  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

This will close #4747 when merged.

@tpn tpn self-assigned this May 20, 2025
@tpn
tpn requested a review from a team as a code owner May 20, 2025 22:55
@tpn tpn added the 3.0 label May 20, 2025
@tpn tpn added this to CCCL May 20, 2025
@github-project-automation github-project-automation Bot moved this to Todo in CCCL May 20, 2025
@cccl-authenticator-app cccl-authenticator-app Bot moved this from Todo to In Review in CCCL May 20, 2025
scan_op: ScanOpType,
initial_value: Any = None,
items_per_thread: int = 1,
items_per_thread: int = 4,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

important: having default number of items per thread is a bit concerning. Even if there was a reasonable default, it's dictated by user data, not CUB. Since we don't know what user is about to pass (yet), I'd probably favor interface that forces user to spell this number explicitly.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel like this PR is the most sensible place to land that. I'll work on another commit now.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tpn
tpn requested a review from gevtushenko May 20, 2025 23:51
@github-actions

Copy link
Copy Markdown
Contributor
🟨 CI finished in 1h 24m: Pass: 83%/12 | Total: 1h 50m | Avg: 9m 12s | Max: 20m 00s
  • 🟨 python: Pass: 83%/12 | Total: 1h 50m | Avg: 9m 12s | Max: 20m 00s

    🚨 jobs: Test cuda.cooperative 🚨
      🟩 Build cuda.cccl    Pass: 100%/2   | Total:  6m 31s | Avg:  3m 15s | Max:  3m 16s
      🟩 Build cuda.cooperative Pass: 100%/2   | Total:  6m 47s | Avg:  3m 23s | Max:  3m 28s
      🟩 Build cuda.parallel Pass: 100%/2   | Total: 16m 35s | Avg:  8m 17s | Max:  8m 23s
      🟩 Test cuda.cccl     Pass: 100%/2   | Total:  9m 04s | Avg:  4m 32s | Max:  4m 43s
      🔥 Test cuda.cooperative Pass:   0%/2   | Total: 37m 32s | Avg: 18m 46s | Max: 20m 00s
      🟩 Test cuda.parallel Pass: 100%/2   | Total: 34m 02s | Avg: 17m 01s | Max: 17m 04s
    🟨 cpu
      🟨 amd64              Pass:  83%/12  | Total:  1h 50m | Avg:  9m 12s | Max: 20m 00s
    🟨 ctk
      🟨 12.9               Pass:  83%/12  | Total:  1h 50m | Avg:  9m 12s | Max: 20m 00s
    🟨 cudacxx
      🟨 nvcc12.9           Pass:  83%/12  | Total:  1h 50m | Avg:  9m 12s | Max: 20m 00s
    🟨 cudacxx_family
      🟨 nvcc               Pass:  83%/12  | Total:  1h 50m | Avg:  9m 12s | Max: 20m 00s
    🟨 cxx
      🟨 GCC13              Pass:  83%/12  | Total:  1h 50m | Avg:  9m 12s | Max: 20m 00s
    🟨 cxx_family
      🟨 GCC                Pass:  83%/12  | Total:  1h 50m | Avg:  9m 12s | Max: 20m 00s
    🟨 gpu
      🟨 rtxa6000           Pass:  83%/12  | Total:  1h 50m | Avg:  9m 12s | Max: 20m 00s
    🟨 py_version
      🟨 3.10               Pass:  83%/6   | Total: 56m 20s | Avg:  9m 23s | Max: 20m 00s
      🟨 3.13               Pass:  83%/6   | Total: 54m 11s | Avg:  9m 01s | Max: 17m 32s
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
+/- python
CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
+/- python
CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 12)

# Runner
6 linux-amd64-cpu16
6 linux-amd64-gpu-rtxa6000-latest-1

tpn added 4 commits May 20, 2025 19:00
The lowering error Numba CUDA was reporting was simply due to it not
knowing how to lower the `initial_value` `Complex()` that was being used
in the tests.

Adding in a `lower_constant` to `Complex` was sufficient to fix the
issue.  There are now no restrictions regarding user-defined types and
items per thread greater than 1.
Now that there are no restrictions to usage when items_per_thread is
greater than 1 (i.e. we fixed the user-defined types issue), it makes
sense to use a more optimal default that should yield better performance
out-of-the-box for most use cases.
@tpn
tpn force-pushed the 4747-fix-items_per_thread-restriction-in-cuda-coop branch from 6e59a37 to bef7ddc Compare May 21, 2025 02:00
@github-actions

Copy link
Copy Markdown
Contributor
🟩 CI finished in 27m 41s: Pass: 100%/12 | Total: 1h 51m | Avg: 9m 15s | Max: 19m 52s
  • 🟩 python: Pass: 100%/12 | Total: 1h 51m | Avg: 9m 15s | Max: 19m 52s

    🟩 cpu
      🟩 amd64              Pass: 100%/12  | Total:  1h 51m | Avg:  9m 15s | Max: 19m 52s
    🟩 ctk
      🟩 12.9               Pass: 100%/12  | Total:  1h 51m | Avg:  9m 15s | Max: 19m 52s
    🟩 cudacxx
      🟩 nvcc12.9           Pass: 100%/12  | Total:  1h 51m | Avg:  9m 15s | Max: 19m 52s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/12  | Total:  1h 51m | Avg:  9m 15s | Max: 19m 52s
    🟩 cxx
      🟩 GCC13              Pass: 100%/12  | Total:  1h 51m | Avg:  9m 15s | Max: 19m 52s
    🟩 cxx_family
      🟩 GCC                Pass: 100%/12  | Total:  1h 51m | Avg:  9m 15s | Max: 19m 52s
    🟩 gpu
      🟩 rtxa6000           Pass: 100%/12  | Total:  1h 51m | Avg:  9m 15s | Max: 19m 52s
    🟩 jobs
      🟩 Build cuda.cccl    Pass: 100%/2   | Total:  6m 56s | Avg:  3m 28s | Max:  3m 35s
      🟩 Build cuda.cooperative Pass: 100%/2   | Total:  6m 56s | Avg:  3m 28s | Max:  3m 35s
      🟩 Build cuda.parallel Pass: 100%/2   | Total: 18m 31s | Avg:  9m 15s | Max:  9m 19s
      🟩 Test cuda.cccl     Pass: 100%/2   | Total:  8m 59s | Avg:  4m 29s | Max:  4m 45s
      🟩 Test cuda.cooperative Pass: 100%/2   | Total: 37m 15s | Avg: 18m 37s | Max: 19m 52s
      🟩 Test cuda.parallel Pass: 100%/2   | Total: 32m 34s | Avg: 16m 17s | Max: 17m 10s
    🟩 py_version
      🟩 3.10               Pass: 100%/6   | Total: 55m 47s | Avg:  9m 17s | Max: 17m 23s
      🟩 3.13               Pass: 100%/6   | Total: 55m 24s | Avg:  9m 14s | Max: 19m 52s
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
+/- python
CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
CUDA Experimental
stdpar
+/- python
CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 12)

# Runner
6 linux-amd64-cpu16
6 linux-amd64-gpu-rtxa6000-latest-1

@tpn
tpn merged commit b30ce9c into NVIDIA:main May 21, 2025
@github-project-automation github-project-automation Bot moved this from In Review to Done in CCCL May 21, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Fix items_per_thread > 1 and user-defined types restriction in cuda.coop.

2 participants