test/docs/mthreads_offline_deployment.md

54 lines
1.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters!

This file contains ambiguous Unicode characters that may be confused with others in your current locale. If your use case is intentional and legitimate, you can safely ignore this warning. Use the Escape button to highlight these characters.

# 摩尔线程 GPU 国产化离线部署指南
推荐使用摩尔线程官方 MUSA vLLM 基础镜像:
```bash
registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519
```
该镜像负责提供 MUSA 运行时、PyTorch、vLLM 与相关内核。本项目的 `Dockerfile.mthreads` 只叠加通用 Python 依赖和业务代码,不在 Dockerfile 内重新安装 vLLM,也不使用 uv 虚拟环境。
## 1. 拉取官方镜像
```bash
docker pull registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519
```
## 2. 生成融合镜像
```bash
./scripts/package_vendor_gpu_image.sh \
--vendor mthreads \
--base-image registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519 \
-v s4000_4.3.5_d0519
```
## 3. 启动服务
```bash
ASR_IMAGE=unis/qwen3-asr:mthreads-s4000_4.3.5_d0519 \
docker compose -f docker-compose-mthreads.yml up -d
```
多卡示例:
```bash
MTHREADS_VISIBLE_DEVICES=0,1 docker compose -f docker-compose-mthreads.yml up -d
```
项目默认同时兼容 `MTHREADS_VISIBLE_DEVICES`、`MUSA_VISIBLE_DEVICES` 和 `CUDA_VISIBLE_DEVICES`。
## 4. 离线交付目录
```bash
./export_offline_bundle.sh \
--type mthreads \
--mthreads-base registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519
```
## 说明
- 容器内设备探测命令使用 `mthreads-gmi`
- `ACCELERATOR=mthreads` 会走 vLLM GPU 路径
- 不建议在项目层重新 `pip install vllm`,避免破坏厂商镜像内的 MUSA 依赖匹配关系