Skip to content

[STF] Ensure we generate CUDA graphs which always have the same topology - #4705

Merged
caugonnet merged 1 commit into
NVIDIA:mainfrom
caugonnet:stf_stable_cuda_graphs
May 16, 2025
Merged

[STF] Ensure we generate CUDA graphs which always have the same topology#4705
caugonnet merged 1 commit into
NVIDIA:mainfrom
caugonnet:stf_stable_cuda_graphs

Conversation

@caugonnet

Copy link
Copy Markdown
Contributor

Some algorithms may reorder the edges in a CUDA graph, resulting in different topologies for the same code, and therefore prevent to update executable graphs.

Description

closes

Checklist

  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

â€Ķlogy so that the update is more likely to succeed
@copy-pr-bot

copy-pr-bot Bot commented May 15, 2025

Copy link
Copy Markdown
Contributor

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@caugonnet caugonnet self-assigned this May 15, 2025
@caugonnet caugonnet added the stf Sequential Task Flow programming model label May 15, 2025
@cccl-authenticator-app cccl-authenticator-app Bot moved this from Todo to In Progress in CCCL May 15, 2025
@caugonnet

Copy link
Copy Markdown
Contributor Author

/ok to test f9fd8a9

@caugonnet

Copy link
Copy Markdown
Contributor Author

/ok to test f9fd8a9

@caugonnet
caugonnet marked this pull request as ready for review May 15, 2025 13:10
@caugonnet
caugonnet requested a review from a team as a code owner May 15, 2025 13:10
@caugonnet
caugonnet requested a review from pciolkosz May 15, 2025 13:10
@cccl-authenticator-app cccl-authenticator-app Bot moved this from In Progress to In Review in CCCL May 15, 2025
@caugonnet
caugonnet enabled auto-merge (squash) May 15, 2025 13:29
@github-actions

Copy link
Copy Markdown
Contributor
ðŸŸĐ CI finished in 41m 29s: Pass: 100%/26 | Total: 6h 14m | Avg: 14m 23s | Max: 28m 52s | Hits: 67%/14642
  • ðŸŸĐ cudax: Pass: 100%/26 | Total: 6h 14m | Avg: 14m 23s | Max: 28m 52s | Hits: 67%/14642

    ðŸŸĐ cpu
      ðŸŸĐ amd64              Pass: 100%/22  | Total:  5h 22m | Avg: 14m 39s | Max: 28m 52s | Hits:  68%/12298 
      ðŸŸĐ arm64              Pass: 100%/4   | Total: 51m 43s | Avg: 12m 55s | Max: 14m 12s | Hits:  61%/2344  
    ðŸŸĐ ctk
      ðŸŸĐ 12.0               Pass: 100%/3   | Total: 39m 48s | Avg: 13m 16s | Max: 14m 27s | Hits:  68%/1463  
      ðŸŸĐ 12.8               Pass: 100%/23  | Total:  5h 34m | Avg: 14m 32s | Max: 28m 52s | Hits:  66%/13179 
    ðŸŸĐ cudacxx
      ðŸŸĐ nvcc12.0           Pass: 100%/3   | Total: 39m 48s | Avg: 13m 16s | Max: 14m 27s | Hits:  68%/1463  
      ðŸŸĐ nvcc12.8           Pass: 100%/23  | Total:  5h 34m | Avg: 14m 32s | Max: 28m 52s | Hits:  66%/13179 
    ðŸŸĐ cudacxx_family
      ðŸŸĐ nvcc               Pass: 100%/26  | Total:  6h 14m | Avg: 14m 23s | Max: 28m 52s | Hits:  67%/14642 
    ðŸŸĐ cxx
      ðŸŸĐ Clang14            Pass: 100%/2   | Total: 26m 18s | Avg: 13m 09s | Max: 13m 37s | Hits:  61%/1176  
      ðŸŸĐ Clang15            Pass: 100%/1   | Total: 14m 26s | Avg: 14m 26s | Max: 14m 26s | Hits:  61%/586   
      ðŸŸĐ Clang16            Pass: 100%/1   | Total: 15m 20s | Avg: 15m 20s | Max: 15m 20s | Hits:  61%/586   
      ðŸŸĐ Clang17            Pass: 100%/1   | Total: 15m 21s | Avg: 15m 21s | Max: 15m 21s | Hits:  61%/586   
      ðŸŸĐ Clang18            Pass: 100%/1   | Total: 14m 40s | Avg: 14m 40s | Max: 14m 40s | Hits:  61%/586   
      ðŸŸĐ Clang19            Pass: 100%/4   | Total: 46m 48s | Avg: 11m 42s | Max: 14m 19s | Hits:  71%/2344  
      ðŸŸĐ GCC10              Pass: 100%/2   | Total: 29m 27s | Avg: 14m 43s | Max: 15m 00s | Hits:  61%/1176  
      ðŸŸĐ GCC11              Pass: 100%/1   | Total: 16m 07s | Avg: 16m 07s | Max: 16m 07s | Hits:  61%/586   
      ðŸŸĐ GCC12              Pass: 100%/1   | Total: 17m 15s | Avg: 17m 15s | Max: 17m 15s | Hits:  61%/586   
      ðŸŸĐ GCC13              Pass: 100%/8   | Total:  1h 34m | Avg: 11m 52s | Max: 15m 27s | Hits:  70%/4688  
      ðŸŸĐ MSVC14.39          Pass: 100%/1   | Total: 12m 40s | Avg: 12m 40s | Max: 12m 40s | Hits:  95%/287   
      ðŸŸĐ MSVC14.42          Pass: 100%/1   | Total: 13m 41s | Avg: 13m 41s | Max: 13m 41s | Hits:  95%/287   
      ðŸŸĐ NVHPC25.3          Pass: 100%/2   | Total: 57m 16s | Avg: 28m 38s | Max: 28m 52s | Hits:  58%/1168  
    ðŸŸĐ cxx_family
      ðŸŸĐ Clang              Pass: 100%/10  | Total:  2h 12m | Avg: 13m 17s | Max: 15m 21s | Hits:  65%/5864  
      ðŸŸĐ GCC                Pass: 100%/12  | Total:  2h 37m | Avg: 13m 08s | Max: 17m 15s | Hits:  67%/7036  
      ðŸŸĐ MSVC               Pass: 100%/2   | Total: 26m 21s | Avg: 13m 10s | Max: 13m 41s | Hits:  95%/574   
      ðŸŸĐ NVHPC              Pass: 100%/2   | Total: 57m 16s | Avg: 28m 38s | Max: 28m 52s | Hits:  58%/1168  
    ðŸŸĐ gpu
      ðŸŸĐ h100               Pass: 100%/2   | Total: 19m 05s | Avg:  9m 32s | Max: 11m 19s | Hits:  80%/1172  
      ðŸŸĐ rtx2080            Pass: 100%/24  | Total:  5h 55m | Avg: 14m 47s | Max: 28m 52s | Hits:  65%/13470 
    ðŸŸĐ jobs
      ðŸŸĐ Build              Pass: 100%/23  | Total:  5h 48m | Avg: 15m 10s | Max: 28m 52s | Hits:  62%/12884 
      ðŸŸĐ Test               Pass: 100%/3   | Total: 25m 24s | Avg:  8m 28s | Max:  9m 49s | Hits:  99%/1758  
    ðŸŸĐ sm
      ðŸŸĐ 90                 Pass: 100%/3   | Total: 30m 22s | Avg: 10m 07s | Max: 11m 19s | Hits:  73%/1758  
      ðŸŸĐ 90a                Pass: 100%/1   | Total: 12m 15s | Avg: 12m 15s | Max: 12m 15s | Hits:  61%/586   
    ðŸŸĐ std
      ðŸŸĐ 17                 Pass: 100%/4   | Total:  1h 04m | Avg: 16m 06s | Max: 28m 24s | Hits:  60%/2342  
      ðŸŸĐ 20                 Pass: 100%/22  | Total:  5h 09m | Avg: 14m 04s | Max: 28m 52s | Hits:  68%/12300 
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
+/- CUDA Experimental
stdpar
python
CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
libcu++
CUB
Thrust
+/- CUDA Experimental
stdpar
python
CCCL C Parallel Library
Catch2Helper

🏃‍ Runner counts (total jobs: 26)

# Runner
17 linux-amd64-cpu16
4 linux-arm64-cpu16
2 windows-amd64-cpu16
2 linux-amd64-gpu-rtx2080-latest-1
1 linux-amd64-gpu-h100-latest-1

::std::unordered_set<cudaGraphNode_t> seen;
::std::vector<cudaGraphNode_t> result;

for (cudaGraphNode_t node : nodes)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nitpick: we should reserve the space

Suggested change
for (cudaGraphNode_t node : nodes)
result.reserve(nodes.size());
for (cudaGraphNode_t node : nodes)

@caugonnet
caugonnet merged commit 0ee684b into NVIDIA:main May 16, 2025
@github-project-automation github-project-automation Bot moved this from In Review to Done in CCCL May 16, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stf Sequential Task Flow programming model

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

2 participants