Skip to content

Measure NFP8 dispatch-view overhead - #6347

Open
q10 wants to merge 1 commit into
pytorch:mainfrom
q10:export-D116893980
Open

q10 wants to merge 1 commit into
pytorch:mainfrom
q10:export-D116893980

Conversation

@q10

@q10 q10 commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

Summary:
Add a benchmark for the ROCm NFP8 fn-to-fnuz dispatch view.

The benchmark measures the isolated Tensor.view(dtype) cost and the end-to-end NFP8 TBE forward time. It keeps performance evidence separate from the correctness diff.

Differential Revision: D116893980

Summary:
Add a benchmark for the ROCm NFP8 `fn`-to-`fnuz` dispatch view.

The benchmark measures the isolated `Tensor.view(dtype)` cost and the end-to-end NFP8 TBE forward time. It keeps performance evidence separate from the correctness diff.

Differential Revision: D116893980
@meta-cla meta-cla Bot added the cla signed label Sep 21, 2026
@meta-codesync

meta-codesync Bot commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

@q10 has exported this pull request. If you are a Meta employee, you can view the originating Diff in D116893980.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant