Summary
A Metal .pte loads under executor_runner built by run_metal_test.sh, but fails in an application that links the ExecuTorch runtime on its own:
E executorch:metal_backend.cpp:352] Failed to load shared library: dlopen(.../<hash>_so_blob<pid>.so, 0x0005):
Library not loaded: /opt/llvm-openmp/lib/libomp.dylib
E executorch:method.cpp:132] Init failed for backend MetalBackend: 0x22
Cause
On macOS inductor's _get_openmp_args always adds -lomp. The libomp it resolves is the one bundled in the PyTorch wheel (torch/lib/libomp.dylib), whose LC_ID_DYLIB is the absolute path /opt/llvm-openmp/lib/libomp.dylib, and that path is recorded as a dependency of the AOTI object embedded in the .pte. The path exists on no end-user machine.
It goes unnoticed in-tree because run_metal_test.sh already rewrites the same reference in executor_runner to @rpath/libomp.dylib with an rpath into torch/lib. The runner therefore loads PyTorch's libomp first, and dyld then satisfies the model's dependency by install name. An app that ships its own libomp (or none) has no image with that install name loaded, so dlopen fails.
The dependency is not used: nm -u on the embedded object shows no omp / kmp symbol for a Metal-delegated MobileNet, as expected for a graph that runs on the GPU.
Possible directions
- Link the AOTI object with
-Wl,-dead_strip_dylibs for the Metal backend, so dylibs the object does not reference are dropped and libomp stays only if a model really needs it. Inductor has no option for extra link flags today; we do it from outside by pointing torch._inductor.config.cpp.cxx at a wrapper compiler that appends the flag on link steps. The resulting object links only libc++ and libSystem and loads in a process with no libomp at all.
- Independently, a check next to
MetalBackend.codesign_so that the compiled object has no dependency outside /usr/lib and /System would turn this from a load-time surprise on someone else's machine into an export-time error.
Happy to send a PR for the check if that direction sounds right; the link-flag part probably belongs in inductor.
Summary
A Metal
.pteloads underexecutor_runnerbuilt byrun_metal_test.sh, but fails in an application that links the ExecuTorch runtime on its own:Cause
On macOS inductor's
_get_openmp_argsalways adds-lomp. The libomp it resolves is the one bundled in the PyTorch wheel (torch/lib/libomp.dylib), whoseLC_ID_DYLIBis the absolute path/opt/llvm-openmp/lib/libomp.dylib, and that path is recorded as a dependency of the AOTI object embedded in the.pte. The path exists on no end-user machine.It goes unnoticed in-tree because
run_metal_test.shalready rewrites the same reference inexecutor_runnerto@rpath/libomp.dylibwith an rpath intotorch/lib. The runner therefore loads PyTorch's libomp first, and dyld then satisfies the model's dependency by install name. An app that ships its own libomp (or none) has no image with that install name loaded, sodlopenfails.The dependency is not used:
nm -uon the embedded object shows noomp/kmpsymbol for a Metal-delegated MobileNet, as expected for a graph that runs on the GPU.Possible directions
-Wl,-dead_strip_dylibsfor the Metal backend, so dylibs the object does not reference are dropped and libomp stays only if a model really needs it. Inductor has no option for extra link flags today; we do it from outside by pointingtorch._inductor.config.cpp.cxxat a wrapper compiler that appends the flag on link steps. The resulting object links onlylibc++andlibSystemand loads in a process with no libomp at all.MetalBackend.codesign_sothat the compiled object has no dependency outside/usr/liband/Systemwould turn this from a load-time surprise on someone else's machine into an export-time error.Happy to send a PR for the check if that direction sounds right; the link-flag part probably belongs in inductor.