Repository navigation
Define GPU validation tests for GPU-enabled drivers #1472
Description
Activity
- added a parent issue
on May 20, 2026 - addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on May 20, 2026 📋 triage-agent
Triage Assessment
Classification: feature-valid
Summary
This is a valid GPU e2e testing feature request. The proposed umbrella GPU validation layout is feasible against the current Rust e2e harness, and it extends existing GPU device-selection coverage with a CUDA workload execution check without changing runtime behavior.
Investigation
- The current GPU e2e target is a single
gpu_device_selectiontest ine2e/rust/Cargo.toml, with feature gating throughe2e-gpu. tasks/test.tomlcurrently wires both Docker and Podman GPU lanes toOPENSHELL_E2E_*_TEST = "gpu_device_selection"; moving to an umbrellagputarget is straightforward, but the implementation should account for the existing Podman task even though Podman CUDA workload coverage is out of scope..github/workflows/e2e-gpu-test.yamlcurrently setsOPENSHELL_E2E_GPU_PROBE_IMAGEand runsmise run --no-deps --skip-deps e2e:docker:gpu; it does not set the new CUDA workload image variable, so initial behavior should clearly log that CUDA execution validation was skipped.e2e/rust/src/harness/sandbox.rsalready captures combined sandbox creation output throughSandboxGuard::create, which is suitable for checking theOPENSHELL_GPU_WORKLOAD_SUCCESSmarker and reporting useful failure output.- The issue body shows
--image, but the current CLI image override foropenshell sandbox createis--from. Unless a CLI alias is intentionally added in a separate scope, the implementation should use--from "$OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE".
Recommendation
Route this as a buildable test feature after human review. Suggested labels:
feature request,topic:testing,test:e2e, andtest:e2e-gpu. Do not addstate:agent-readyduring triage; that remains a human approval step before implementation.- The current GPU e2e target is a single
- addedtest:e2eRequires end-to-end coverageRequires end-to-end coveragetest:e2e-gpuRequires GPU end-to-end coverageRequires GPU end-to-end coverageand removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on May 20, 2026 - changed the title
[-]Define basic workload tests for GPU-enabled drivers[/-][+]Define GPU validation tests for GPU-enabled drivers[/+]on May 20, 2026 🏗️ build-plan
Implementation Plan
Issue type:
test
Complexity: Medium
Confidence: High — the e2e harness already captures sandbox output; full positive-path CUDA validation depends on having a suitable local/published validation image.Summary
Add an umbrella Rust GPU e2e test target that preserves existing device-selection coverage and adds Docker-focused CUDA workload validation. The CUDA workload test uses a prebuilt GPU validation image directly as the sandbox image through
--from, runs the image default command, and assertsOPENSHELL_GPU_WORKLOAD_SUCCESS.CUDA execution in GPU CI is out of scope for this issue and is tracked by #1487. This issue must provide the harness, local positive-path validation when an image is available, and visible skip behavior when
OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGEis unset.Scope
e2e/rust/Cargo.toml: replace thegpu_device_selectionintegration target with an umbrellagputarget gated bye2e-gpu.e2e/rust/tests/gpu.rs: add the integration-test root for GPU validation modules.e2e/rust/tests/gpu/device_selection.rs: move the existinggpu_device_selection.rstests here without behavior changes.e2e/rust/tests/gpu/workloads.rs: add the CUDA workload validation test.tasks/test.toml: update Docker and Podman GPU e2e task names fromgpu_device_selectiontogpu; Podman workload validation remains out of scope and may skip.e2e/rust/e2e-docker.sh: emit a high-level note when Docker GPU e2e is running withoutOPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE.- Documentation: update
e2e/gpu/README.mdif test(e2e): add GPU workload image artifacts #1484 has landed. If not, add minimal documentation near the Rust e2e/GPU test area or task docs and reconcile after test(e2e): add GPU workload image artifacts #1484 lands.
Implementation Steps
-
Rename the Rust GPU test target from
gpu_device_selectiontogpu. -
Move existing GPU device-selection tests to
e2e/rust/tests/gpu/device_selection.rswithout changing assertions. -
Add
e2e/rust/tests/gpu.rsthat includes the GPU validation modules. -
Add
e2e/rust/tests/gpu/workloads.rsfor CUDA workload validation. -
Read
OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE, trimming whitespace and treating empty values as unset. -
If unset, print:
skipping CUDA GPU workload validation: OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE is not setthen return without failing.
-
If set, call:
openshell sandbox create \ --gpu \ --from "$OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE"with no explicit command override, so the image default command runs.
-
Assert combined captured output contains:
OPENSHELL_GPU_WORKLOAD_SUCCESS -
Include the image name and captured output in failure messages.
-
Update Docker and Podman GPU task wiring to use
--test gpu. -
Update docs with the env var, marker contract, local run command, skip behavior, and relationship to test(e2e): add GPU workload image artifacts #1484/test(ci): enable CUDA GPU execution validation in GPU CI #1487.
Test Plan
-
Unit tests: Preserve existing parser/helper tests from
gpu_device_selection.rsafter moving them under the newgputarget. -
Integration tests: Run:
cargo test --manifest-path e2e/rust/Cargo.toml --features e2e-gpu --test gpu -- --list -
E2E tests:
- Run
mise run e2e:docker:gpuwithoutOPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE; verify device-selection tests run and CUDA workload validation logs the explicit skip. - If test(e2e): add GPU workload image artifacts #1484 images are available locally, set
OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGEto the localcuda-basicimage and rerunmise run e2e:docker:gpu; verify marker-positive CUDA workload validation passes. - Podman CUDA workload validation is not required for this issue.
- Run
Risks & Open Questions
- test(e2e): add GPU workload image artifacts #1484 is not a hard dependency for starting implementation. It provides the intended local image artifacts for positive-path validation, but Define GPU validation tests for GPU-enabled drivers #1472 can land the harness and skip behavior independently.
- test(ci): enable CUDA GPU execution validation in GPU CI #1487 tracks enabling CUDA workload validation in GPU CI once a published immutable image reference exists.
- Rust’s default test harness has no dynamic skipped status, so missing CUDA image configuration may appear as
ok; the test and runner must log the skip clearly. - LSM compatibility risk is low. This change does not add host process inspection,
/proc/<pid>/exechecks, or new host-side binary execution.
Documentation Impact
Update e2e/GPU documentation to describe:
- the umbrella
gputest target, OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE,OPENSHELL_GPU_WORKLOAD_SUCCESS,- skip behavior when the image env var is unset,
- local validation with test(e2e): add GPU workload image artifacts #1484 image artifacts when available,
- CI enforcement deferred to test(ci): enable CUDA GPU execution validation in GPU CI #1487.
Revision 1 — initial plan
- addedstate:agent-readyApproved for agent implementationApproved for agent implementationand removedstate:review-readyReady for human reviewReady for human review
on Jun 3, 2026 Follow-up finding from the workload manifest PR:
The
smoke-passandsmoke-failworkload images are useful for validating the workload harness itself, but they are not inherently GPU workloads. They exercise manifest parsing, sandbox command execution, success-marker detection, failure-marker detection, and expected-failure handling.For the current GPU validation PR, it still makes sense to run them in the GPU suite so the manifest runner is covered alongside
cuda-basic. Longer term, we should consider generalizing workload validation so non-GPU smoke workloads can run in a generic e2e workload suite, while GPU-specific workloads are selected viarequirements.gpuor similar manifest metadata.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Problem Statement
Instead of only running
nvidia-smiin a sandbox, OpenShell should have GPU validation tests that exercise basic GPU functionality from inside a GPU-enabled sandbox.The existing GPU e2e tests cover device discovery, selection, and visibility. This issue adds an execution-focused GPU validation test that verifies a sandbox can run a basic CUDA workload. The test structure should leave room for future validation classes such as OpenCL, Vulkan, and additional driver integrations.
Proposed Design
Use “GPU validation tests” as the umbrella term. Device-selection tests are one class of GPU validation; CUDA execution tests are another.
Initial Scope
The first implementation should target the Docker compute driver only. Docker is currently the most mature GPU-enabled e2e path.
The test code should be organized so additional GPU-enabled drivers can reuse the same validation categories later, but Podman, Kubernetes, and VM integration are out of scope for this issue. Podman should be the first follow-up driver.
Test Layout
Move the existing GPU device-selection tests into a broader GPU test target and add the CUDA execution test alongside them:
Update the e2e Cargo test target to use the umbrella GPU test:
The existing device-selection test behavior should remain unchanged after the move.
CUDA Execution Test
The initial CUDA execution test should cover the default driver behavior when
--gpuis requested.The test should create a Docker-backed OpenShell sandbox with GPU enabled and use the configured CUDA workload image directly as the sandbox image:
openshell sandbox create \ --gpu \ --from "$OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE"The workload image should run its default command. Do not add a workload command override in the first implementation.
The first PR should not add
--gpu-deviceor per-device workload permutations. Existing GPU e2e tests continue to cover device selection andnvidia-smivisibility behavior.Workload Image Contract
Add
OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGEfor the CUDA execution image.The image must be usable directly as an OpenShell sandbox image. It should run without requiring nested container tooling or network access inside the sandbox.
The image base is not part of the contract. Workload images are separate test artifacts with their own dependency sets and command contracts.
The workload image must emit this stable success marker to stdout or stderr when the GPU workload completes successfully:
The e2e test should require both:
OPENSHELL_GPU_WORKLOAD_SUCCESS.On failure, test output should include the workload image name, sandbox context, exit status when available, stdout, and stderr.
Missing Image Behavior
If
OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGEis unset or empty, the CUDA execution test may skip itself.Rust’s default test harness has no dynamic skipped status, so this may appear as
okin the Cargo summary. The test must still emit a clear log message, for example:The Docker GPU e2e runner should also print a high-level note when the variable is unset so CI/local logs clearly show that CUDA execution coverage did not run.
Out of Scope
This issue should not add host/runtime preflight checks for the CUDA workload image. Existing GPU setup and visibility checks remain separate concerns. The new execution test should focus on whether an OpenShell sandbox can run a GPU workload successfully once the GPU-enabled driver and workload image are available.
This issue should not define or publish the reference workload images. That work is tracked separately in #1476.
Follow-ups
OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGEand require CUDA execution coverage.Acceptance Criteria
OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGEthroughopenshell sandbox create --from.OPENSHELL_GPU_WORKLOAD_SUCCESSin stdout or stderr.OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGEis reported clearly in test/e2e output and does not fail the lane initially.gputest target.