Repository navigation
Conversation
A run-only library test that runs one Pantheon workload on the GPUs of a node: memory bandwidth, a March C- memory test, a memory retention test or an FP16 compute load. The test passes if every selected GPU completed the workload, Pantheon's verification found no errors, and the card reported no uncorrectable errors. The score and the peak temperature of each GPU are performance metrics. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Saqib Khan <11544614+saqibkh@users.noreply.github.com>
|
A limitation I found after opening this, which reviewers should know about. Pantheon 1.2.2 does not honour The test is not affected when the job owns all the GPUs of the node, which is what it asks for: it sets I reproduced it on a two-GPU machine with |
|
Thanks @saqibkh for your PR. We are not actively developing the testlib inside ReFrame at the moment. We are considering though reviving it in a separate repository, in which case we will let you know. |
This adds a GPU diagnostics test to the test library, next to
gpu_burn_check.What it does
The test runs one workload of Pantheon, an open-source (Apache-2.0) GPU diagnostics suite for NVIDIA and AMD cards, on the GPUs of a node. It is parameterized over four workloads:
memory_readmarch_testmemory_retentiontensor_virusgpu_burn_checkanswers whether a card survives a sustained load. This test answers whether each part of the card gives correct results, and how fast, so a degraded card shows up either as a failed sanity check or as a performance regression against the reference of its node type.The sanity check passes if every selected GPU completed the workload, Pantheon's own verification of the results found no errors, and the card reported no uncorrectable errors during the run. It fails if Pantheon found no GPU and fell back to its CPU backend, so a pass always means hardware was tested. The score and the peak temperature of each GPU are reported as performance metrics.
What it needs
The
pantheonexecutable in thePATH(pip install pantheon-gpu) and the compiler of the GPU toolchain,nvccorhipcc. Pantheon compiles its workloads on the node, for the card it finds there, the first time it runs. The test is run-only and has no sources in this repository.How I tested it
With this branch (4.11.0-dev4) and with 4.10.4:
-S platform=mock), which needs no GPU and is an easy way to try the test.memory_readandmarch_testpass on an NVIDIA RTX 3060 with CUDA, and report the score and the peak temperature.flake8is clean with the project's settings, and the docs build with Sphinx 8.2.3 without warnings from the new module.I have not run it on an AMD card or through a batch scheduler.
Disclosure
I maintain Pantheon. If a test that depends on an external tool does not belong in the library, I am happy to keep it in our repository instead and would appreciate a pointer to where such tests are best listed. The test was written with the help of an AI assistant, and I ran and checked it as described above.
🤖 Generated with Claude Code