Skip to content

feat: add hardware platform and plugin support - #1565

Open
zhangts20 wants to merge 2 commits into
mainfrom
hardware-platform-plugin
Open

zhangts20 wants to merge 2 commits into
mainfrom
hardware-platform-plugin

Conversation

@zhangts20

@zhangts20 zhangts20 commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

简介

把硬件抽象和第三方插件拆成两层,方便后续其他硬件以及 pip 的注册,当前提交不涉及现有推理路径。对外入口是 get_hardware_backend,当前还未被调用,启动服务不会涉及本地提交代码。

lightllm/platform/
├── __init__.py
├── ops
│   ├── __init__.py
│   ├── act.py
│   ├── norm.py
│   └── sampling.py
├── backends
│   ├── __init__.py
│   ├── cuda_like
│   │   ├── __init__.py
│   │   ├── graph.py
│   │   ├── runtime.py
│   │   └── ops
│   │       ├── __init__.py
│   │       ├── act.py
│   │       ├── norm.py
│   │       └── sampling.py
│   └── ascend
│       └── ...
├── base
│   ├── __init__.py
│   ├── backend.py
│   ├── graph.py
│   ├── registry.py
│   └── runtime.py
└── plugin
    ├── __init__.py
    ├── att.py
    ├── common.py
    └── ops.py

backends 下面放置具体的硬件后端实现,包括 graph、runtime 和 ops;platform/ops 只定义各后端共用的算子(当前是部分,仅包含 norm、act 和 sampling),不是具体的某个硬件实现。base 是所有基类以及 HardwareBackend 的注册;plugin 实现 pip-installed 包的注册;__init__.py 提供对外接口 get_hardware_backend,用来初始化全局实例(后续接入主推理流程时,可在确定具体硬件后端时立即调用以提早暴露错误)。

做法

  • 硬件 backend:为避免使用硬件 A 时受到硬件 B 部分代码的影响,采用懒加载的方式。register_platform 时只登记类(HardwareBackend 子类)的路径,后续调用 get_hardware_backend 时才真正初始化。
  • Plugin 加载:pip entry point -> CLI 匹配 -> register 返回 modules -> import module -> 写入注册表。
  • 算子组装:backend 初始化时以内置 ops 为默认值,再用已登记的 pip 字段替换。未登记的字段保持内置实现。
  • CLI:--extra_ops / --extra_att(逗号分隔,对应 entry point 的左边名字)。

入口只有 get_hardware_backend():先 configure_plugins(),再按 hardware_platform 构造 backend。

用法

[build-system]
requires = ["setuptools>=61"]
build-backend = "setuptools.build_meta"

[project]
name = "lightllm-example-plugin"
version = "0.0.1"
description = "Example LightLLM ops and att plugins"
requires-python = ">=3.10"

[project.entry-points."lightllm.ops_plugin"]
example_ops = "lightllm_example_plugin.register:register_ops"

[project.entry-points."lightllm.att_plugin"]
example_att = "lightllm_example_plugin.register:register_att"

[tool.setuptools.packages.find]
where = ["."]
include = ["lightllm_example_plugin*"]

同时注册 lightllm.ops_plugin 和 lightllm.att_plugin。过程如下:

  1. pip install xxx 安装包
  2. ops 的 modules 由 lightllm_example_plugin.register:register_ops 函数返回
  3. att 的 modules 由 lightllm_example_plugin.register:register_att 函数返回
  4. 在 ops 的 module 中调用 register_ops 登记算子字段,在 att 的 module 中使用 register_att_backend 装饰器修饰类
  5. 配合启动参数 --extra_ops example_ops 和 --extra_att example_att 完成注册并使用

新增 backend 和算子

和算子注册不同,这部分需修改 lightllm/platform 代码。

  • 新增 backend
  1. 在 backends 下新增目录,存放 runtime、graph 和 ops。
  2. 写一个 HardwareBackend 子类,类上标明 platform_name,构造时传入自己的 runtime、graph 和 ops 或者使用已有定义。
  3. 调用 register_platform 注册类路径,参考上述说明,实际使用时检查正确性。

如果实现和 CUDA 一样,可以不新建目录,像沐曦类似继承 CudaLikeBackend,只改 platform_name 并登记类的路径。最后,为启动参数 --hardware_platform 增加启动选项。

对于新增的 backend,可以部分继承 CUDA 实现,但只能按照 runtime、graph 和 ops 来组合。如 runtime 和 graph 分别使用 CudaLikeRuntime 和 CudaLikeGraph;算子可以整体使用 CUDA_LIKE_OPS,也可以部分使用如 norm 用 CUDA_LIKE_NORM_OPS、act 和 sampling 等用自己定义的。

  • 新增算子

在 platform/ops 新增 dataclasses 类,再给 PlatformOps 新增字段。每个平台的 PlatformOps 初始化都要带上这个新增算子。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant