Skip to content

[Feature Request] Add RPC Backend / Remote Offload Support #181

Description

@alcoftTAO

Is your feature request related to a problem? Please describe.

Currently, llama-cpp-python exposes several compute backends (CPU, CUDA, Vulkan, SYCL, Metal), but there is no first-class support for llama.cpp's RPC backend. This makes it difficult to:

  • Offload tensor computations to a remote GPU machine.
  • Share a single GPU across multiple Python processes or containers without each one holding the full model in VRAM.
  • Run inference on machines with limited local compute (e.g., a small CPU-only node) while leveraging a powerful remote accelerator.

At the moment, users who need RPC must either build llama.cpp from source with -DGGML_RPC=ON and manage the server binary externally, or resort to fragile workarounds. A native Python API for this would be a huge quality-of-life improvement.

Describe the solution you'd like

It would be wonderful if llama-cpp-python could expose the RPC backend through its existing tensor_split / main_gpu parameters, for example:

from llama_cpp import Llama

llm = Llama(
    model_path="model.gguf",
    n_gpu_layers=-1,
    main_gpu=0,
    rpc_servers=["192.168.1.50:42232"],  # <= new kwarg
    tensor_split=[2, 1]  # <= 0 could be this machine, 1 could be the first remote server
)

Additionally, a small helper to launch the RPC server from Python would be ideal:

from llama_cpp import start_rpc_server

server = start_rpc_server(
    host="0.0.0.0",
    port=42232,
    backend="cuda",   # or "vulkan", "cpu", ...
)
# ... run inference on the client side ...
server.stop()

Even a thin wrapper around ggml-rpc-server started via subprocess would be a great start.

Describe alternatives you've considered

  • Manual build + external server: Compile llama.cpp with -DGGML_RPC=ON, run ./ggml-rpc-server & by hand, and point the client at it. Would work, but it's fragile and hard to document for new users.

Additional context

llama.cpp already has a mature RPC backend in C++ that speaks a simple TCP protocol.

Happy to help test if that would be helpful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions