Skip to content
 
 

Latest commit

 

History

879 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

GDRCopy for Moore Threads

This is a GDRCopy project ported for Moore Threads GPU platform.

Table of Contents

.
├── src/                  # Source code directory
│   ├── gdrdrv/          # Kernel driver module
│   └── *.c              # User-space API implementation
├── include/              # Header files directory
├── tests/                # Test programs directory
├── packages/             # Package configuration files directory
├── scripts/              # Build and installation scripts directory
├── build/                # Build output directory
└── Makefile             # Build script

Introduction

While GPUDirect RDMA is meant for direct access to GPU memory from third-party devices, it is possible to use these same APIs to create perfectly valid CPU mappings of the GPU memory.

The advantage of a CPU driven copy is the very small overhead involved. That might be useful when low latencies are required.

What is inside

GDRCopy offers the infrastructure to create user-space mappings of GPU memory, which can then be manipulated as if it was plain host memory (caveats apply here).

A simple by-product of it is a copy library with the following characteristics:

  • An initial memory pinning phase is required, which is potentially expensive,

  • Fast H-D, because of write-combining.

  • Slow D-H, because the GPU BAR, which backs the mappings, can't be prefetched and so burst reads transactions are not generated through PCIE

The library comes with a few tests like:

  • mtgdrcopy_sanity, which contains unit tests for the library and the driver.
  • mtgdrcopy_copybw, a minimal application which calculates the R/W bandwidth for a specific buffer size.
  • mtgdrcopy_copylat, a benchmark application which calculates the R/W copy latency for a range of buffer sizes.
  • mtgdrcopy_apiperf, an application for benchmarking the latency of each GDRCopy API call.

GDRCopy

GDRCopy is a low-latency GPU memory copy library based on NVIDIA GDRCopy driver, designed for Moore Threads MTGPU GPUDirect RDMA technology. It enables CPU to directly map and access GPU memory for low-latency data transfers.

Version Information

Current version components:

  • mtgdrdrv kernel module: 0.1.0
  • libgdrapi user-space library: 1.0.0

Compatibility Strategy

libgdrapi 1.x.x versions are designed to be backward compatible with mtgdrdrv 0.x.x kernel module versions. This means users can update the user-space library independently without upgrading the kernel module.

Build Options

Basic Build Options

  • make all # Build all components
  • make driver # Build kernel driver module
  • make lib # Build user-space API library
  • make exes # Build test programs
  • make clean # Clean build artifacts

Installation Options

  • make install # Install all components
  • make lib_install # Install user-space API library
  • make exes_install # Install test programs
  • make drv_install # Install kernel driver module

Package Build Options

  • make build # Build directory structure for packaging
  • make deb # Create single DEB package (traditional way)
  • make deb-bundle # Create single DEB package with all components (bundle way)

Split Package Build Options (New Feature)

  • make build-kmod # Build kernel module package structure
  • make build-lib # Build API library package structure
  • make build-tests # Build test tools package structure
  • make build-meta # Build meta package structure
  • make deb-kmod # Create kernel module DEB package
  • make deb-lib # Create API library DEB package
  • make deb-tests # Create test tools DEB package
  • make deb-meta # Create meta package DEB package
  • make deb-all # Create all DEB packages

Usage

Traditional Way (Single Package)

make build
make deb

Bundle Way (Single Package with All Components)

make build
make deb-bundle

New Way (Split Packages)

# Build each package separately
make build-kmod
make build-lib
make build-tests
make build-meta

# Or build all packages at once
make deb-all

Package Description

After splitting, there will be 4 packages:

  1. mtgdrdrv-dkms - Kernel module package, containing DKMS kernel module and related configurations
  2. libgdrapi - API library package, containing user-space API library and header files
  3. mtgdrcopy-tests - Test tools package, containing test tools and related dependencies
  4. mtgdrcopy - Meta package, depending on the above three packages for backward compatibility

Testing

To verify package building, generate all packages:

make deb-all

This creates the four .deb files in the build/ directory.

Requirements

GDRCopy has been tested on S80, S4000 and S5000.

For CPU compatibility, GDRCopy has primarily been tested on Intel, AMD and Hygon platforms. Functional compatibility does not imply identical performance. Performance can vary with the GPU model, PCIe topology, CPU cache policy, IOMMU configuration and the instruction set selected for the CPU.

ARM builds are expected to compile successfully, but compatibility across different ARM platforms has not been validated and ARM packages are not currently provided. ARM support is experimental. Other architectures supported by the upstream NVIDIA GDRCopy project are outside the supported scope of this project.

MUSA DRIVER is required in all deployment environments. For host-based use, install the complete MUSA DRIVER stack on the host before building or installing GDRCopy. This includes the GPU driver, the driver development headers, and the compatible MUSA toolkit required by the build and tests.

For container-based use, install the mtgdrdrv-dkms package on the host, and install the libgdrapi and mtgdrcopy-tests DEB packages inside the container. Ensure that the container has the required permissions and that the GPU devices and driver interfaces are mapped correctly into the container.

We recommend using the official MUSA containers and following the official container workflow. For details, see the MUSA container documentation and the official container image directory.

Root privileges are necessary to load or install the kernel-mode device driver on the host.

Build and installation

from source

$ make
$ sudo ./insmod.sh
$ export LD_LIBRARY_PATH=$(readlink -f ./src):${LD_LIBRARY_PATH}

Tests

After making from source successfully, you can run tests

$ cd tests
$ ./mtgdrcopy_copybw

All tests should be able to compile and run on latest MUSA platform. For old version of MUSA, some subtests of mtgdrcopy_sanity will fail. The performance varies by CPU and GPU.

Restrictions and known issues

GDRCopy works with regular MUSA device memory only, as returned by musaMalloc. In particular, it does not work with MUSA managed memory.

gdr_pin_buffer() accepts any addresses returned by musaMalloc and its family. In contrast, gdr_map() requires that the pinned address is aligned to the GPU page. Neither MUSA Runtime nor Driver APIs guarantees that GPU memory allocation functions return aligned addresses. Users are responsible for proper alignment of addresses passed to the library. This is a restriction from the original NVIDIA driver. Unaligned address may be supported by further updates.

Multiple musaMalloc allocations may appear contiguous in the GPU virtual address space. However, passing an address range that spans multiple allocations to gdr_pin_buffer() or gdr_map() is not generally guaranteed to work. Support depends on the MUSA DRIVER version and hardware platform, and applications should not rely on this behavior.

Bug filing

For reporting issues, please open an issue in the public repository: https://github.com/MooreThreads/mtgdrcpy/issues.

Acknowledgment

The project is maintained by Moore Threads and is distributed through the public repository. Please review the platform requirements and known limitations before use.

If you find this software useful in your work, please cite: R. Shi et al., "Designing efficient small message transfer mechanism for inter-node MPI communication on InfiniBand GPU clusters," 2014 21st International Conference on High Performance Computing (HiPC), Dona Paula, 2014, pp. 1-10, doi: 10.1109/HiPC.2014.7116873.

About

A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages