Is your feature request related to a problem? Please describe.
Currently, llama-cpp-python exposes several compute backends (CPU, CUDA, Vulkan, SYCL, Metal), but there is no first-class support for llama.cpp's RPC backend. This makes it difficult to:
- Offload tensor computations to a remote GPU machine.
- Share a single GPU across multiple Python processes or containers without each one holding the full model in VRAM.
- Run inference on machines with limited local compute (e.g., a small CPU-only node) while leveraging a powerful remote accelerator.
At the moment, users who need RPC must either build llama.cpp from source with -DGGML_RPC=ON and manage the server binary externally, or resort to fragile workarounds. A native Python API for this would be a huge quality-of-life improvement.
Describe the solution you'd like
It would be wonderful if llama-cpp-python could expose the RPC backend through its existing tensor_split / main_gpu parameters, for example:
from llama_cpp import Llama
llm = Llama(
model_path="model.gguf",
n_gpu_layers=-1,
main_gpu=0,
rpc_servers=["192.168.1.50:42232"], # <= new kwarg
tensor_split=[2, 1] # <= 0 could be this machine, 1 could be the first remote server
)
Additionally, a small helper to launch the RPC server from Python would be ideal:
from llama_cpp import start_rpc_server
server = start_rpc_server(
host="0.0.0.0",
port=42232,
backend="cuda", # or "vulkan", "cpu", ...
)
# ... run inference on the client side ...
server.stop()
Even a thin wrapper around ggml-rpc-server started via subprocess would be a great start.
Describe alternatives you've considered
- Manual build + external server: Compile
llama.cpp with -DGGML_RPC=ON, run ./ggml-rpc-server & by hand, and point the client at it. Would work, but it's fragile and hard to document for new users.
Additional context
llama.cpp already has a mature RPC backend in C++ that speaks a simple TCP protocol.
Happy to help test if that would be helpful.
Is your feature request related to a problem? Please describe.
Currently,
llama-cpp-pythonexposes several compute backends (CPU, CUDA, Vulkan, SYCL, Metal), but there is no first-class support for llama.cpp's RPC backend. This makes it difficult to:At the moment, users who need RPC must either build
llama.cppfrom source with-DGGML_RPC=ONand manage the server binary externally, or resort to fragile workarounds. A native Python API for this would be a huge quality-of-life improvement.Describe the solution you'd like
It would be wonderful if
llama-cpp-pythoncould expose the RPC backend through its existingtensor_split/main_gpuparameters, for example:Additionally, a small helper to launch the RPC server from Python would be ideal:
Even a thin wrapper around
ggml-rpc-serverstarted viasubprocesswould be a great start.Describe alternatives you've considered
llama.cppwith-DGGML_RPC=ON, run./ggml-rpc-server &by hand, and point the client at it. Would work, but it's fragile and hard to document for new users.Additional context
llama.cpp already has a mature RPC backend in C++ that speaks a simple TCP protocol.
Happy to help test if that would be helpful.