Skip to content

{2025.06}[2025a] PantheonGPU 1.2.2 w/ CUDA 12.8.0 - #1738

Open
saqibkh wants to merge 1 commit into
EESSI:mainfrom
saqibkh:pantheongpu-2025.06
Open

saqibkh wants to merge 1 commit into
EESSI:mainfrom
saqibkh:pantheongpu-2025.06

Conversation

@saqibkh

@saqibkh saqibkh commented Oct 1, 2026

Copy link
Copy Markdown

Adds PantheonGPU 1.2.2 (CUDA 12.8.0, gfbf/2025a) to the 2025.06 stack, next to ollama in accel/nvidia.

Pantheon is an open-source (Apache-2.0) GPU diagnostics suite: each workload loads one part of a card, verifies the results to detect silent data corruption, and reads the card's error counters before and after the run. A user on a node that mounts EESSI would run, for example, pantheon --test march_test --duration 60 --gpu all after module load PantheonGPU.

The easyconfig was merged into easybuild-easyconfigs on 2026-09-30 (easybuilders/easybuild-easyconfigs#26992, tagged for 5.4.1) and is referenced by its merge commit until that release. A maintainer built it on their cluster and ran the installed module on an H200 and a P100.

Notes for the build:

  • All dependencies are modules of the 2025a stack (Python-bundle-PyPI, SciPy-bundle, pyNVML) plus CUDA/12.8.0 from accel/nvidia.
  • The sanity check runs a workload on Pantheon's CPU backend, so it passes on a build node without a GPU. On a node with a GPU, Pantheon compiles its workloads with nvcc on the first run, into the user's cache directory, so it needs the full CUDA toolkit from host_injections at run time, as other CUDA software in EESSI does.
  • Disclosure: we maintain Pantheon.

🤖 Generated with Claude Code

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@bedroge

bedroge commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

Thanks a lot for your (first) contribution! The PR looks really good already, two small remarks/issues:

so it needs the full CUDA toolkit from host_injections at run time, as other CUDA software in EESSI does

This one is a little bit different, though. We only ship runtime libraries in EESSI, as the CUDA EULA does not allow redistribution of e.g. the compilers. For almost all software, only the drivers have to be made available on client systems (i.e. not the full CUDA toolkit), and then things just work, as the runtime libraries are available. PantheonGPU needs the compilers, though, so it does need the full toolkit to be available. This is not a blocker, but just something to be aware of for end-users of the software.

All dependencies are modules of the 2025a stack (Python-bundle-PyPI, SciPy-bundle, pyNVML) plus CUDA/12.8.0 from accel/nvidia.

For GPU builds, all CPU-only dependencies need to be in place already. It looks like that's the case for everything except pyNVML. So that would still have to be added to the stack before we can trigger builds for this. Could you please open a separate PR for this and add it to https://github.com/EESSI/software-layer/blob/main/easystacks/software.eessi.io/2025.06/eessi-2025.06-eb-5.4.0-2025a.yml?

@saqibkh

saqibkh commented Oct 1, 2026 •

Copy link
Copy Markdown
Author

Thanks for the quick look. The pyNVML entry is in #1740. On the toolkit: understood. We had read host_injections as the usual route for software that compiles at run time, and we will say in Pantheon documentation that an EESSI site needs the full CUDA SDK installed there, not only the drivers.

@bedroge

bedroge commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

Small correction/addition to my previous message: we probably do need to modify something in how we handle CUDA dependencies. Currently every CUDA dependency is dropped to a build dependency, meaning that the CUDA module won't be loaded at runtime; we rely on RPATH to find the CUDA runtime libraries. Obviously, that won't work for PantheonGPU, as it needs to be able to find the compilers. So we probably need to make an exception in EESSI's EasyBuild hooks file for this, we'll look into that.

@ocaisa

ocaisa commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

@saqibkh looking at the readme for PantheonGPU, we could actually compile all the kernels in advance with make PLATFORM=CUDA, we may just need to tweak the EasyBuild recipe to hard set the correct cc capability. In EESSI we already have separate installation paths for the spectrum of compute capabilities and we detect the right one at runtime.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

2025.06-software.eessi.io 2025.06 version of software.eessi.io accel:nvidia

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants