Enable hybrid CUDA IPC and EFA GPU exchange - #413
Draft
devavret wants to merge 1 commit into
Draft
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
smandcuda_ipcalongside EFA/SRD so UCX can use PCIe P2P for reachable neighboring GPUs and SRD for the remaining pairsUCX_MAX_RNDV_RAILSto 2 because it was the better of the two hybrid settings in the complete-suite A/BThis is a draft stacked on #412. The UCX patch should be removed once the upstream change is available in the UCX release used by the worker image.
Validation
All query tests used local-NVMe SF3K, 8 GPU workers, 2 drivers, 3 iterations, and no exchange compression. Streaming aggregation was enabled only as an experimental config change and is not part of this PR.
MAX(total_revenue)equality instability: pure SRD and rails 1 returned no row, while rails 2 returned one row.This remains a draft because the transport-selection mechanism works, but the full-suite data does not yet demonstrate an end-to-end performance win.