Skip to content

Define GPU validation tests for GPU-enabled drivers #1472

Description

@linear

Problem Statement

Instead of only running nvidia-smi in a sandbox, OpenShell should have GPU validation tests that exercise basic GPU functionality from inside a GPU-enabled sandbox.

The existing GPU e2e tests cover device discovery, selection, and visibility. This issue adds an execution-focused GPU validation test that verifies a sandbox can run a basic CUDA workload. The test structure should leave room for future validation classes such as OpenCL, Vulkan, and additional driver integrations.

Proposed Design

Use “GPU validation tests” as the umbrella term. Device-selection tests are one class of GPU validation; CUDA execution tests are another.

Initial Scope

The first implementation should target the Docker compute driver only. Docker is currently the most mature GPU-enabled e2e path.

The test code should be organized so additional GPU-enabled drivers can reuse the same validation categories later, but Podman, Kubernetes, and VM integration are out of scope for this issue. Podman should be the first follow-up driver.

Test Layout

Move the existing GPU device-selection tests into a broader GPU test target and add the CUDA execution test alongside them:

e2e/rust/tests/gpu.rs
e2e/rust/tests/gpu/
  device_selection.rs
  execution.rs
  cuda.rs

Update the e2e Cargo test target to use the umbrella GPU test:

[[test]]
name = "gpu"
path = "tests/gpu.rs"
required-features = ["e2e-gpu"]

The existing device-selection test behavior should remain unchanged after the move.

CUDA Execution Test

The initial CUDA execution test should cover the default driver behavior when --gpu is requested.

The test should create a Docker-backed OpenShell sandbox with GPU enabled and use the configured CUDA workload image directly as the sandbox image:

openshell sandbox create \
  --gpu \
  --from "$OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE"

The workload image should run its default command. Do not add a workload command override in the first implementation.

The first PR should not add --gpu-device or per-device workload permutations. Existing GPU e2e tests continue to cover device selection and nvidia-smi visibility behavior.

Workload Image Contract

Add OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE for the CUDA execution image.

The image must be usable directly as an OpenShell sandbox image. It should run without requiring nested container tooling or network access inside the sandbox.

The image base is not part of the contract. Workload images are separate test artifacts with their own dependency sets and command contracts.

The workload image must emit this stable success marker to stdout or stderr when the GPU workload completes successfully:

OPENSHELL_GPU_WORKLOAD_SUCCESS

The e2e test should require both:

  • the sandbox workload command exits successfully,
  • combined stdout/stderr contains OPENSHELL_GPU_WORKLOAD_SUCCESS.

On failure, test output should include the workload image name, sandbox context, exit status when available, stdout, and stderr.

Missing Image Behavior

If OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE is unset or empty, the CUDA execution test may skip itself.

Rust’s default test harness has no dynamic skipped status, so this may appear as ok in the Cargo summary. The test must still emit a clear log message, for example:

skipping CUDA GPU execution test: OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE is not set

The Docker GPU e2e runner should also print a high-level note when the variable is unset so CI/local logs clearly show that CUDA execution coverage did not run.

Out of Scope

This issue should not add host/runtime preflight checks for the CUDA workload image. Existing GPU setup and visibility checks remain separate concerns. The new execution test should focus on whether an OpenShell sandbox can run a GPU workload successfully once the GPU-enabled driver and workload image are available.

This issue should not define or publish the reference workload images. That work is tracked separately in #1476.

Follow-ups

  • Add Podman integration for the GPU validation tests.
  • Define and publish reference GPU validation images in test(e2e): define GPU validation image artifacts #1476.
  • Once a reference CUDA image is published, update GPU CI to set OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE and require CUDA execution coverage.
  • Add additional validation classes such as OpenCL or Vulkan when suitable workload images exist.
  • Consider refactoring shared GPU validation helpers when a second workload or driver integration is added.

Acceptance Criteria

  • Existing GPU device-selection tests are moved under the new GPU validation test layout without behavior changes.
  • A new CUDA execution validation test is added for Docker-backed GPU sandboxes.
  • The CUDA execution test uses OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE through openshell sandbox create --from.
  • The CUDA execution test runs the image default command.
  • The CUDA execution test asserts OPENSHELL_GPU_WORKLOAD_SUCCESS in stdout or stderr.
  • Missing OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE is reported clearly in test/e2e output and does not fail the lane initially.
  • The Docker GPU e2e path runs the umbrella gpu test target.
  • Documentation or e2e README content explains how to configure and run the CUDA execution validation locally.

Activity

  1. elezar commented on May 20, 2026

    @elezar
    Member

    📋 triage-agent

    Triage Assessment

    Classification: feature-valid

    Summary

    This is a valid GPU e2e testing feature request. The proposed umbrella GPU validation layout is feasible against the current Rust e2e harness, and it extends existing GPU device-selection coverage with a CUDA workload execution check without changing runtime behavior.

    Investigation

    • The current GPU e2e target is a single gpu_device_selection test in e2e/rust/Cargo.toml, with feature gating through e2e-gpu.
    • tasks/test.toml currently wires both Docker and Podman GPU lanes to OPENSHELL_E2E_*_TEST = "gpu_device_selection"; moving to an umbrella gpu target is straightforward, but the implementation should account for the existing Podman task even though Podman CUDA workload coverage is out of scope.
    • .github/workflows/e2e-gpu-test.yaml currently sets OPENSHELL_E2E_GPU_PROBE_IMAGE and runs mise run --no-deps --skip-deps e2e:docker:gpu; it does not set the new CUDA workload image variable, so initial behavior should clearly log that CUDA execution validation was skipped.
    • e2e/rust/src/harness/sandbox.rs already captures combined sandbox creation output through SandboxGuard::create, which is suitable for checking the OPENSHELL_GPU_WORKLOAD_SUCCESS marker and reporting useful failure output.
    • The issue body shows --image, but the current CLI image override for openshell sandbox create is --from. Unless a CLI alias is intentionally added in a separate scope, the implementation should use --from "$OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE".

    Recommendation

    Route this as a buildable test feature after human review. Suggested labels: feature request, topic:testing, test:e2e, and test:e2e-gpu. Do not add state:agent-ready during triage; that remains a human approval step before implementation.

  2. added
    test:e2eRequires end-to-end coverage
    test:e2e-gpuRequires GPU end-to-end coverage
    and removed
    state:triage-neededOpened without agent diagnostics and needs triage
    on May 20, 2026
  3. changed the title [-]Define basic workload tests for GPU-enabled drivers[/-] [+]Define GPU validation tests for GPU-enabled drivers[/+] on May 20, 2026
  4. elezar commented on May 20, 2026

    @elezar
    Member

    🏗️ build-plan

    Implementation Plan

    Issue type: test
    Complexity: Medium
    Confidence: High — the e2e harness already captures sandbox output; full positive-path CUDA validation depends on having a suitable local/published validation image.

    Summary

    Add an umbrella Rust GPU e2e test target that preserves existing device-selection coverage and adds Docker-focused CUDA workload validation. The CUDA workload test uses a prebuilt GPU validation image directly as the sandbox image through --from, runs the image default command, and asserts OPENSHELL_GPU_WORKLOAD_SUCCESS.

    CUDA execution in GPU CI is out of scope for this issue and is tracked by #1487. This issue must provide the harness, local positive-path validation when an image is available, and visible skip behavior when OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE is unset.

    Scope

    • e2e/rust/Cargo.toml: replace the gpu_device_selection integration target with an umbrella gpu target gated by e2e-gpu.
    • e2e/rust/tests/gpu.rs: add the integration-test root for GPU validation modules.
    • e2e/rust/tests/gpu/device_selection.rs: move the existing gpu_device_selection.rs tests here without behavior changes.
    • e2e/rust/tests/gpu/workloads.rs: add the CUDA workload validation test.
    • tasks/test.toml: update Docker and Podman GPU e2e task names from gpu_device_selection to gpu; Podman workload validation remains out of scope and may skip.
    • e2e/rust/e2e-docker.sh: emit a high-level note when Docker GPU e2e is running without OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE.
    • Documentation: update e2e/gpu/README.md if test(e2e): add GPU workload image artifacts #1484 has landed. If not, add minimal documentation near the Rust e2e/GPU test area or task docs and reconcile after test(e2e): add GPU workload image artifacts #1484 lands.

    Implementation Steps

    1. Rename the Rust GPU test target from gpu_device_selection to gpu.

    2. Move existing GPU device-selection tests to e2e/rust/tests/gpu/device_selection.rs without changing assertions.

    3. Add e2e/rust/tests/gpu.rs that includes the GPU validation modules.

    4. Add e2e/rust/tests/gpu/workloads.rs for CUDA workload validation.

    5. Read OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE, trimming whitespace and treating empty values as unset.

    6. If unset, print:

      skipping CUDA GPU workload validation: OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE is not set
      

      then return without failing.

    7. If set, call:

      openshell sandbox create \
        --gpu \
        --from "$OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE"

      with no explicit command override, so the image default command runs.

    8. Assert combined captured output contains:

      OPENSHELL_GPU_WORKLOAD_SUCCESS
      
    9. Include the image name and captured output in failure messages.

    10. Update Docker and Podman GPU task wiring to use --test gpu.

    11. Update docs with the env var, marker contract, local run command, skip behavior, and relationship to test(e2e): add GPU workload image artifacts #1484/test(ci): enable CUDA GPU execution validation in GPU CI #1487.

    Test Plan

    • Unit tests: Preserve existing parser/helper tests from gpu_device_selection.rs after moving them under the new gpu target.

    • Integration tests: Run:

      cargo test --manifest-path e2e/rust/Cargo.toml --features e2e-gpu --test gpu -- --list
    • E2E tests:

      • Run mise run e2e:docker:gpu without OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE; verify device-selection tests run and CUDA workload validation logs the explicit skip.
      • If test(e2e): add GPU workload image artifacts #1484 images are available locally, set OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE to the local cuda-basic image and rerun mise run e2e:docker:gpu; verify marker-positive CUDA workload validation passes.
      • Podman CUDA workload validation is not required for this issue.

    Risks & Open Questions

    Documentation Impact

    Update e2e/GPU documentation to describe:


    Revision 1 — initial plan

  5. added and removed on Jun 3, 2026
  6. moved this from Todo to In progress in OpenShell Roadmapon Jun 16, 2026
  7. elezar commented on Jun 17, 2026

    @elezar
    Member

    Follow-up finding from the workload manifest PR:

    The smoke-pass and smoke-fail workload images are useful for validating the workload harness itself, but they are not inherently GPU workloads. They exercise manifest parsing, sandbox command execution, success-marker detection, failure-marker detection, and expected-failure handling.

    For the current GPU validation PR, it still makes sense to run them in the GPU suite so the manifest runner is covered alongside cuda-basic. Longer term, we should consider generalizing workload validation so non-GPU smoke workloads can run in a generic e2e workload suite, while GPU-specific workloads are selected via requirements.gpu or similar manifest metadata.

  8. moved this from In progress to Done in OpenShell Roadmapon Jun 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions