Skip to content

GPU: Apple Metal backend, off by default - #15859

Merged
ktf merged 1 commit into
AliceO2Group:devfrom
ktf:pr15859
Sep 29, 2026
Merged

ktf merged 1 commit into
AliceO2Group:devfrom
ktf:pr15859

Conversation

@ktf

@ktf ktf commented Sep 28, 2026

Copy link
Copy Markdown
Member

The backend itself: the Objective-C++ host side, the .metal kernel source and
its build rules, plus the CMake to enable them.

macOS had no usable GPU backend before this. It ships OpenCL 1.2, below the 2.x
the OpenCL backend requires, so find_package(OpenCL) there could never produce
one; the version check dropped it again a few lines later. That lookup is now
skipped on Apple, which leaves CUDA_ENABLED, OPENCL_ENABLED and HIP_ENABLED all
necessarily off, so the Darwin arm of the backend dispatch was dead code and
goes with it. Metal takes its place.

The build rules also give Metal its entry in the no-fast-math table, so that once
GPUCA_DETERMINISTIC_MODE reaches GPUCA_DETERMINISTIC_MODE_MAP_NO_FAST_MATH it
drops its fast math flags like the other backends rather than keeping whatever
the build type gave it.

Off unless asked for. FindO2GPU.cmake leaves ENABLE_METAL=OFF on macOS and the
subdirectory is gated on METAL_ENABLED, so macOS keeps running on the CPU until
the whole chain is validated.

Apple toolchain only: the source goes .metal -> AIR through xcrun metal and
nothing else, with no SPIR-V translation step in between.

Requires -std=metal4.1, the first MSL version with a generic address space.
Earlier versions reject an unannotated pointer with 'pointer type must have
explicit address space qualifier' and an unannotated 'this' with 'cannot
initialize object parameter', both of which GPUCommonDefAPI.h relies on for
GPUgeneric() and GPUdDefault(). Verified against Xcode 27, which ships
metal4.1; Xcode 26 and earlier stop at metal4.0.

The Metal frameworks ship with every macOS, so finding them says nothing about
whether the backend can be built; the configure compiles a three-line kernel as
MSL 4.1 to answer that directly. The deployment target has no say either: -std=
is what picks the target OS, and MACOSX_DEPLOYMENT_TARGET and
-mmacosx-version-min are both ignored by the Metal compiler. AUTO therefore
turns Metal off on an older toolchain instead of failing somewhere in the middle
of the build, and an explicit ENABLE_METAL=ON says why it cannot be honoured.

@ktf
ktf requested review from a team and davidrohr as code owners September 28, 2026 17:59

@davidrohr davidrohr left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In general, nothing scary for me. In any case, it will be compiled only for MacOS.
I didn't read it in detail enough to understand anything, and I have no way of testing it anyway.
I saw that a comment still mentions OpenCL, so perhaps one should go through the code once and do some cleanup. I guess a lot was copy and pasted from the OpenCL host code.

int32_t GPUReconstructionMetal::InitDevice_Runtime()
{
// Propagate processing settings to PoCL runtime.
// Won't affect other OpenCL runtimes.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Might need some cleanup, this is for metal, not OpenCL

@ktf

ktf commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

Yes, this is indeed my old copy pastes from last year plus further corrections from my friends for metal 4.1. I will ask for a cleanup of this commit.

The backend itself: the Objective-C++ host side, the .metal kernel source and
its build rules, plus the CMake to enable them.

macOS had no usable GPU backend before this. It ships OpenCL 1.2, below the 2.x
the OpenCL backend requires, so find_package(OpenCL) there could never produce
one; the version check dropped it again a few lines later. That lookup is now
skipped on Apple, which leaves CUDA_ENABLED, OPENCL_ENABLED and HIP_ENABLED all
necessarily off, so the Darwin arm of the backend dispatch was dead code and
goes with it. Metal takes its place.

The build rules also give Metal its entry in the no-fast-math table, so that once
GPUCA_DETERMINISTIC_MODE reaches GPUCA_DETERMINISTIC_MODE_MAP_NO_FAST_MATH it
drops its fast math flags like the other backends rather than keeping whatever
the build type gave it.

Off unless asked for. FindO2GPU.cmake leaves ENABLE_METAL=OFF on macOS and the
subdirectory is gated on METAL_ENABLED, so macOS keeps running on the CPU until
the whole chain is validated.

Apple toolchain only: the source goes .metal -> AIR through xcrun metal and
nothing else, with no SPIR-V translation step in between.

Requires -std=metal4.1, the first MSL version with a generic address space.
Earlier versions reject an unannotated pointer with 'pointer type must have
explicit address space qualifier' and an unannotated 'this' with 'cannot
initialize object parameter', both of which GPUCommonDefAPI.h relies on for
GPUgeneric() and GPUdDefault(). Verified against Xcode 27, which ships
metal4.1; Xcode 26 and earlier stop at metal4.0.

The Metal frameworks ship with every macOS, so finding them says nothing about
whether the backend can be built; the configure compiles a three-line kernel as
MSL 4.1 to answer that directly. The deployment target has no say either: -std=
is what picks the target OS, and MACOSX_DEPLOYMENT_TARGET and
-mmacosx-version-min are both ignored by the Metal compiler. AUTO therefore
turns Metal off on an older toolchain instead of failing somewhere in the middle
of the build, and an explicit ENABLE_METAL=ON says why it cannot be honoured.
@ktf
ktf merged commit 93e2a54 into AliceO2Group:dev Sep 29, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants