Skip to content

restore TBE backward host-side launch config - #6326

Closed
q10 wants to merge 1 commit into
pytorch:mainfrom
q10:export-D120626878
Closed

q10 wants to merge 1 commit into
pytorch:mainfrom
q10:export-D120626878

Conversation

@q10

@q10 q10 commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Summary:
Restores the ROCm V1 TBE backward launch configuration that was dropped while fixing the host-only CUDA build: max_segment_length_per_warp returns to 16384, and BT_block_size follows the active GPU warp size.

Moves kWarpSize and kWarpSizeHost() from cuda_prelude.cuh into the host-safe utils/warp_size.h, so generated host sources can use the runtime query without including device intrinsics.

Marks the internal-linkage kWarpSize declarations [[maybe_unused]] because GCC diagnoses them in generated host translation units where only kWarpSizeHost() is used.

Defers the ROCm runtime query until dev_weights is on GPU. CPU dispatches retain the prior fallback of 64 and do not initialize the HIP runtime.

This follows up D119363782 and #6278, corresponding to #6323.

Reviewed By: gchalump

Differential Revision: D120626878

@meta-cla meta-cla Bot added the cla signed label Sep 18, 2026
@meta-codesync

meta-codesync Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

@q10 has exported this pull request. If you are a Meta employee, you can view the originating Diff in D120626878.

Summary:
Restores the ROCm V1 TBE backward launch configuration that was dropped while fixing the host-only CUDA build: `max_segment_length_per_warp` returns to 16384, and `BT_block_size` follows the active GPU warp size.

Moves `kWarpSize` and `kWarpSizeHost()` from `cuda_prelude.cuh` into the host-safe `utils/warp_size.h`, so generated host sources can use the runtime query without including device intrinsics.

Marks the internal-linkage `kWarpSize` declarations `[[maybe_unused]]` because GCC diagnoses them in generated host translation units where only `kWarpSizeHost()` is used.

Defers the ROCm runtime query until `dev_weights` is on GPU. CPU dispatches retain the prior fallback of 64 and do not initialize the HIP runtime.

This follows up D119363782 and pytorch#6278, corresponding to pytorch#6323.

Reviewed By: gchalump

Differential Revision: D120626878
@meta-codesync meta-codesync Bot changed the title restore TBE backward host-side launch config (#6323) restore TBE backward host-side launch config Sep 18, 2026
@q10
q10 force-pushed the export-D120626878 branch from 9a02faf to 40af862 Compare September 18, 2026 18:52
@meta-codesync meta-codesync Bot closed this in 659e41f Sep 18, 2026
@meta-codesync meta-codesync Bot added the Merged label Sep 18, 2026
@meta-codesync

meta-codesync Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

This pull request has been merged in 659e41f.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant