Skip to content

[testlib] Add a GPU diagnostics test based on Pantheon - #3738

Closed
saqibkh wants to merge 1 commit into
reframe-hpc:developfrom
saqibkh:testlib/pantheon-gpu-diagnostics
Closed

saqibkh wants to merge 1 commit into
reframe-hpc:developfrom
saqibkh:testlib/pantheon-gpu-diagnostics

Conversation

@saqibkh

@saqibkh saqibkh commented Sep 28, 2026

Copy link
Copy Markdown

This adds a GPU diagnostics test to the test library, next to gpu_burn_check.

What it does

The test runs one workload of Pantheon, an open-source (Apache-2.0) GPU diagnostics suite for NVIDIA and AMD cards, on the GPUs of a node. It is parameterized over four workloads:

Workload What it loads
memory_read memory bandwidth
march_test memory cells, with a March C- pattern
memory_retention memory cells holding their charge
tensor_virus the FP16 compute path

gpu_burn_check answers whether a card survives a sustained load. This test answers whether each part of the card gives correct results, and how fast, so a degraded card shows up either as a failed sanity check or as a performance regression against the reference of its node type.

The sanity check passes if every selected GPU completed the workload, Pantheon's own verification of the results found no errors, and the card reported no uncorrectable errors during the run. It fails if Pantheon found no GPU and fell back to its CPU backend, so a pass always means hardware was tested. The score and the peak temperature of each GPU are reported as performance metrics.

What it needs

The pantheon executable in the PATH (pip install pantheon-gpu) and the compiler of the GPU toolchain, nvcc or hipcc. Pantheon compiles its workloads on the node, for the card it finds there, the first time it runs. The test is run-only and has no sources in this repository.

How I tested it

With this branch (4.11.0-dev4) and with 4.10.4:

  • All four workloads pass on Pantheon's CPU backend (-S platform=mock), which needs no GPU and is an easy way to try the test.
  • memory_read and march_test pass on an NVIDIA RTX 3060 with CUDA, and report the score and the peak temperature.
  • flake8 is clean with the project's settings, and the docs build with Sphinx 8.2.3 without warnings from the new module.

I have not run it on an AMD card or through a batch scheduler.

reframe -c hpctestlib/microbenchmarks/gpu/pantheon.py -S platform=mock -S duration=3 \
        -S valid_systems='*' -S valid_prog_environs='*' -r

Disclosure

I maintain Pantheon. If a test that depends on an external tool does not belong in the library, I am happy to keep it in our repository instead and would appreciate a pointer to where such tests are best listed. The test was written with the help of an AI assistant, and I ran and checked it as described above.

🤖 Generated with Claude Code

A run-only library test that runs one Pantheon workload on the GPUs of a
node: memory bandwidth, a March C- memory test, a memory retention test or
an FP16 compute load. The test passes if every selected GPU completed the
workload, Pantheon's verification found no errors, and the card reported
no uncorrectable errors. The score and the peak temperature of each GPU
are performance metrics.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Saqib Khan <11544614+saqibkh@users.noreply.github.com>
@saqibkh

saqibkh commented Sep 28, 2026

Copy link
Copy Markdown
Author

A limitation I found after opening this, which reviewers should know about.

Pantheon 1.2.2 does not honour CUDA_VISIBLE_DEVICES. It enumerates the GPUs of the node through NVML and starts the workload on each of them. If a job is given only some of the node's GPUs through that variable, without device cgroups hiding the rest from NVML, the workload fails on the hidden cards with invalid device ordinal and this test reports those GPUs as failed.

The test is not affected when the job owns all the GPUs of the node, which is what it asks for: it sets exclusive_access and takes num_gpus_per_node from the partition's device configuration. It is also not affected when devices lists only the GPUs the job was given and their numbering matches the node's.

I reproduced it on a two-GPU machine with CUDA_VISIBLE_DEVICES=0. The fix belongs in Pantheon, and I will note here when a release carries it.

@vkarak

vkarak commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks @saqibkh for your PR. We are not actively developing the testlib inside ReFrame at the moment. We are considering though reviving it in a separate repository, in which case we will let you know.

@vkarak vkarak closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants