|
|
||
|---|---|---|
| .pi-crg-mcp | ||
| app | ||
| crg-mcp-plugin | ||
| docs | ||
| environments | ||
| scripts | ||
| test_web | ||
| tests | ||
| vendor/qwenasr | ||
| .dockerignore | ||
| .env.example | ||
| .gitignore | ||
| Dockerfile.cpu | ||
| Dockerfile.gpu | ||
| Dockerfile.iluvatar | ||
| Dockerfile.metax | ||
| Dockerfile.mthreads | ||
| README.md | ||
| REALTIME_ASR_TROUBLESHOOTING.md | ||
| docker-compose-cpu.yml | ||
| docker-compose-iluvatar.yml | ||
| docker-compose-metax.yml | ||
| docker-compose-mthreads.yml | ||
| docker-compose.yml | ||
| export_offline_bundle.sh | ||
| pyproject.toml | ||
| pyrightconfig.json | ||
| start.py | ||
| test_diarization.py | ||
| test_speed.py | ||
| uv.lock | ||
| 曦云系列_通用GPU_mx-smi使用手册_CN_V14.pdf | ||
README.md
Qwen3-ASR
Ready-to-use Local Speech Recognition API Service
Speech recognition API service centered on Qwen3-ASR, with NVIDIA CUDA vLLM, MetaX/MuXi MACA vLLM, and CPU Rust backends, OpenAI API compatibility, Alibaba Cloud Speech API compatibility, and a Paraformer realtime websocket capability.
Live Demo Site
- Web Demo: https://asr.vect.one
Demo
Release 1.0.1
v1.0.1is the current patch release.v1.0.0introduced a large breaking refactor relative to the earliermainbranch. If you are upgrading frommain, read the release notes before reusing old deployment assumptions.Key breaking changes:
- Python dependency management is now
uv-based (pyproject.toml+uv.lock);requirements*.txtare gone- Runtime stack changed to
NVIDIA/MetaX GPU -> vLLM,CPU/macOS -> vendored QwenASR RustMLX/ Apple Silicon GPU path has been removed;mpsis normalized tocpu- macOS / Apple Silicon now defaults to
qwen3-asr-0.6b; setQWEN3_ASR_MODELto override itENABLED_MODELShas been removed
Features
- Hybrid Runtime Stack - Uses auto-selected Qwen3-ASR for offline inference and Paraformer realtime for websocket streaming
- Speaker Diarization - Automatic multi-speaker identification using CAM++ model
- OpenAI API Compatible - Supports
/v1/audio/transcriptionsendpoint, works with OpenAI SDK - Alibaba Cloud API Compatible - Supports Alibaba Cloud Speech RESTful API and WebSocket streaming protocol
- WebSocket Streaming - Real-time streaming speech recognition with low latency
- Smart Far-Field Filtering - Automatically filters far-field sounds and ambient noise in streaming ASR
- Intelligent Audio Segmentation - VAD-based greedy merge algorithm for automatic long audio splitting
- GPU Batch Processing - Batch inference support, 2-3x faster than sequential processing
- Resource-Aware Runtime - Auto-selects the appropriate Qwen3-ASR model for the current machine
Acknowledgements
- Qwen3-ASR provides the official model family and multimodal/vLLM usage guidance
- QwenASR provides the CPU Rust backend vendored by this project
Quick Deployment
1. Docker Deployment (Recommended)
# Copy and edit configuration
cp .env.example .env
# Edit .env to set API_KEY (optional)
# Compose defaults:
# /opt/dep/asr/models -> /app/models
# /opt/dep/asr/data -> /app/data
# /opt/dep/asr/data/logs, temp, tasks live under this data mount
# Optional: override any host mount root in .env
# export MODEL_STORAGE_DIR=/data/qwen3-asr-models
# export DATA_STORAGE_DIR=/data/qwen3-asr-data
# Start service (NVIDIA GPU version)
docker-compose up -d
# Or MetaX/MuXi GPU version
docker-compose -f docker-compose-metax.yml up -d
# Or Iluvatar/Tianshu GPU version
docker-compose -f docker-compose-iluvatar.yml up -d
# Or Moore Threads / MUSA GPU version
docker-compose -f docker-compose-mthreads.yml up -d
# Or CPU version
docker-compose -f docker-compose-cpu.yml up -d
# NVIDIA multi-GPU auto mode (one instance per visible GPU)
CUDA_VISIBLE_DEVICES=0,1,2,3 docker-compose up -d
# MetaX/MuXi multi-GPU auto mode
METAX_VISIBLE_DEVICES=0,1 docker-compose -f docker-compose-metax.yml up -d
# Iluvatar/Tianshu multi-GPU auto mode
ILUVATAR_VISIBLE_DEVICES=0,1 docker-compose -f docker-compose-iluvatar.yml up -d
# Moore Threads / MUSA multi-GPU auto mode
MTHREADS_VISIBLE_DEVICES=0,1 docker-compose -f docker-compose-mthreads.yml up -d
Service URLs:
- API Endpoint:
http://localhost:17003 - API Docs:
http://localhost:17003/docs
Optional built-in rate limit settings:
NGINX_RATE_LIMIT_RPS(global requests/sec,0= disabled)NGINX_RATE_LIMIT_BURST(global burst,0= auto use RPS)
docker run (alternative):
# NVIDIA GPU version
docker run -d --name qwen3-asr \
--gpus all \
-p 17003:8000 \
-e ACCELERATOR=nvidia \
-e CUDA_VISIBLE_DEVICES=0,1,2,3 \
-e API_KEY=your_api_key \
-v /opt/dep/asr/models:/app/models \
-v /opt/dep/asr/data:/app/data \
unis/qwen3-asr:gpu-latest
# MetaX/MuXi GPU version
docker run -d --name qwen3-asr-metax \
--privileged \
--network=host \
--pid=host \
--ipc=host \
-v /dev:/dev \
-v /opt/mxdriver:/opt/mxdriver:ro \
-e ACCELERATOR=metax \
-e PORT=17003 \
-e METAX_VISIBLE_DEVICES=0 \
-v /opt/dep/asr/models:/app/models \
-v /opt/dep/asr/data:/app/data \
unis/qwen3-asr:metax-latest
# CPU version
docker run -d --name qwen3-asr \
-p 17003:8000 \
-v /opt/dep/asr/models:/app/models \
-v /opt/dep/asr/data:/app/data \
unis/qwen3-asr:cpu-latest
Note: NVIDIA GPU images default to CUDA 13.0/cu130 with
torch 2.11.0+vllm 0.20.0. Developers can rebuildDockerfile.gpufor CUDA 12.6, CUDA 13.0, or another backend by overriding Docker build args. MetaX/MuXi images useDockerfile.metaxon top of an official MetaX vLLM image. In field deployments, use host networking plus privileged/devand/opt/mxdrivermounts so bothmx-smiand the MetaX PyTorch runtime can initialize devices. CPU images now supportqwen3-asr-0.6bvia the bundled QwenASR Rust backend. The default CPU image uses a portable Rust target; setQWENASR_RUST_TARGET_CPU=nativeonly for self-built, host-specific images. On CUDA vLLM and CPU Rust,word_timestamps=truenow triggers the forced aligner automatically. On macOS / Apple Silicon, Qwen3-ASR now runs through the Rust CPU backend.start.pynow forces the vLLM multiprocessing method tospawnso startup does not hit CUDA re-initialization failures in forked subprocesses.
Custom GPU backend builds:
# Default GPU build: CUDA 13.0 / PyTorch cu130
docker build -t qwen3-asr:gpu-cu130 -f Dockerfile.gpu .
# CUDA 12.6 build for older deployments
docker build -t qwen3-asr:gpu-cu126 -f Dockerfile.gpu \
--build-arg PYTORCH_BASE_IMAGE=pytorch/pytorch:2.11.0-cuda12.6-cudnn9-runtime \
--build-arg PYTORCH_CUDA_INDEX=https://download.pytorch.org/whl/cu126 \
--build-arg CUDA_NVCC_PACKAGE=cuda-nvcc-12-6 \
--build-arg TORCH_CUDA_ARCH_LIST="8.0;8.6;8.9" \
.
# CUDA 13.0 build when your driver/toolchain requires it
docker build -t qwen3-asr:gpu-cu130 -f Dockerfile.gpu \
--build-arg PYTORCH_BASE_IMAGE=pytorch/pytorch:2.11.0-cuda13.0-cudnn9-runtime \
--build-arg PYTORCH_CUDA_INDEX=https://download.pytorch.org/whl/cu130 \
--build-arg CUDA_NVCC_PACKAGE=cuda-nvcc-13-0 \
--build-arg TORCH_CUDA_ARCH_LIST="12.0+PTX" \
.
# MetaX/MuXi build: fuse this project into an official MetaX vLLM image
./scripts/package_vendor_gpu_image.sh \
--vendor metax \
--base-image <official-metax-vllm-image> \
-v n260-3.7.0.38
# Iluvatar/Tianshu build: fuse this project into the official Iluvatar vLLM image
docker pull registry.iluvatar.com.cn:10443/customer/sz/vllm0.17.0-4.4.0-x86:v5
./scripts/package_vendor_gpu_image.sh \
--vendor iluvatar \
--base-image registry.iluvatar.com.cn:10443/customer/sz/vllm0.17.0-4.4.0-x86:v5 \
-v vllm0.17.0-4.4.0-v5
# Moore Threads / MUSA build: fuse this project into the official MUSA vLLM image
docker pull registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519
./scripts/package_vendor_gpu_image.sh \
--vendor mthreads \
--base-image registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519 \
-v s4000_4.3.5_d0519
For MetaX/MuXi offline delivery, see docs/metax_offline_deployment.md. For Iluvatar/Tianshu offline delivery, see docs/iluvatar_offline_deployment.md. For Moore Threads / MUSA offline delivery, see docs/mthreads_offline_deployment.md.
Offline Deployment: You can now build a timestamped offline delivery folder that includes the image archive, compose file, env template, host-dir init script, and usage docs. The export script uses plain docker build + docker save, so it does not depend on buildx:
# 1. Build an offline delivery folder
./export_offline_bundle.sh --type gpu
# or
./export_offline_bundle.sh --type cpu
# or MetaX/MuXi GPU
./export_offline_bundle.sh \
--type metax \
--metax-base cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.17.0-maca.ai3.5.3.307-torch2.8-py312-ubuntu22.04-amd64 \
--skip-models
# or Iluvatar/Tianshu GPU
./export_offline_bundle.sh --type iluvatar --iluvatar-base registry.iluvatar.com.cn:10443/customer/sz/vllm0.17.0-4.4.0-x86:v5
# or Moore Threads / MUSA GPU
./export_offline_bundle.sh --type mthreads --mthreads-base registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519
# or build both in one bundle
./export_offline_bundle.sh --type all
# 2. Prepare models separately, without deleting existing model files
./scripts/download-models.sh --models-dir /opt/dep/asr/models
# 3. Copy the generated folder to the offline server
scp -r build-file/<timestamp>-all user@server:/opt/dep/asr/
# 4. On the offline server
cd /opt/dep/asr/<timestamp>-all
./init_host_dirs.sh
gunzip -c qwen3-asr-gpu-<timestamp>-amd64.tar.gz | docker load
gunzip -c qwen3-asr-cpu-<timestamp>-amd64.tar.gz | docker load
# NVIDIA GPU
docker compose up -d
# or MetaX/MuXi GPU
# docker compose -f docker-compose-metax.yml up -d
# or Iluvatar/Tianshu GPU
# docker compose -f docker-compose-iluvatar.yml up -d
# or Moore Threads / MUSA GPU
# docker compose -f docker-compose-mthreads.yml up -d
# or CPU
# docker compose -f docker-compose-cpu.yml up -d
Detailed deployment instructions: Deployment Guide
Local Development
System Requirements:
- Python 3.10+
- CUDA 13.0+ for the default GPU image; CUDA 12.6 / 13.0 can be built with Docker args
- FFmpeg (audio format conversion)
Installation:
Runtime dependency locks now default to the GPU stack at the repo root, with CPU kept as a specialized environment:
| Mode | Command | Notes |
|---|---|---|
| NVIDIA GPU (default) | uv sync or ./scripts/sync_gpu_env.sh |
Syncs the root pyproject.toml and uv.lock into .venv, including CUDA 13.0/cu130 torch 2.11.0 / torchaudio 2.11.0 / torchvision 0.26.0 / vllm 0.20.0 |
| MetaX/MuXi GPU | ./scripts/sync_metax_env.sh |
Syncs common dependencies from environments/metax/pyproject.toml; optional GPU-stack install uses the MetaX MACA PyPI index with --no-deps by default |
| Iluvatar/Tianshu GPU | ./scripts/sync_iluvatar_env.sh |
Syncs common dependencies from environments/iluvatar/pyproject.toml; GPU stack should come from the official Iluvatar vLLM image |
| Moore Threads / MUSA GPU | ./scripts/sync_mthreads_env.sh |
Syncs common dependencies from environments/mthreads/pyproject.toml; GPU stack should come from the official Moore Threads MUSA vLLM image |
| CPU (specialized) | ./scripts/sync_cpu_env.sh |
Syncs the dedicated CPU lock in environments/cpu/pyproject.toml into .venv |
| Auto | ./scripts/sync_accel_env.sh |
Chooses MetaX when mx-smi is present, Iluvatar when ixsmi is present, Moore Threads when mthreads-gmi is present, otherwise NVIDIA when nvidia-smi is present, otherwise CPU |
# Clone project
cd qwen3-asr
# Install dependencies (Linux/NVIDIA CUDA)
uv sync
# Start service
source .venv/bin/activate
python start.py
MetaX/MuXi local development:
./scripts/sync_metax_env.sh
source .venv/bin/activate
ACCELERATOR=metax python start.py
Local model storage defaults to ./models under the project root:
./models/
Qwen/
iic/
damo/
Override it when needed:
export MODELS_DIR=/data/qwen3-asr-models
export MODELSCOPE_CACHE=/data
export MODELSCOPE_PATH=$MODELS_DIR
macOS / Apple Silicon local development:
./scripts/sync_cpu_env.sh
source .venv/bin/activate
python start.py
Runtime Defaults
Current runtime behavior on the mainline codebase:
ACCELERATOR=autoresolves tometaxwhenmx-smireports devices, theniluvatarwhenixsmireports devices, thenmthreadswhenmthreads-gmireports devices, otherwisenvidiawhen NVIDIA CUDA is available, otherwisecpuDEVICE=autoresolves to the active accelerator device (cuda:0for NVIDIA/MetaX/Iluvatar GPU, otherwisecpu)DEVICE=mpsis normalized tocpuLinux + NVIDIA CUDAuses officialvLLMLinux + MetaX/MuXi MACAuses the MetaX-compatible PyTorch/vLLM stackLinux + Iluvatar/Tianshuuses the Iluvatar official vLLM image stackLinux + CPUuses vendoredQwenASRRustmacOS / Apple Siliconalso uses vendoredQwenASRRust- macOS / Apple Silicon defaults to
qwen3-asr-0.6b qwen3-asr-1.7bon macOS is only used whenQWEN3_ASR_MODEL=qwen3-asr-1.7bword_timestamps=trueworks on the current offline CUDA and CPU Rust paths- WebSocket streaming does not currently return word-level timestamps
- CAM++ speaker diarization remains required and still follows
DEVICE; on CPU its main hotspot is speaker verification embedding
API Endpoints
OpenAI Compatible API
| Endpoint | Method | Function |
|---|---|---|
/v1/audio/transcriptions |
POST | Audio transcription (OpenAI compatible) |
/v1/models |
GET | Offline model list |
Request Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
file |
file | Preferred when provided | Audio/video file |
audio_address |
string | Optional | Audio/video URL (HTTP/HTTPS), file://, or server-local path. Ignored when file is also provided |
language |
string | Auto-detect | Language code (zh/en/ja) |
enable_speaker_diarization |
bool | true |
Enable speaker diarization |
enable_speaker_identification |
bool | true |
Match registered speaker database when diarization is enabled |
enable_text_cleanup |
bool | true |
Enable text deduplication, boundary-overlap trimming, and filler cleanup |
word_timestamps |
bool | false |
Return word-level timestamps when the backend supports them. Qwen CUDA vLLM and CPU Rust automatically use the forced aligner when enabled. |
hotwords |
string | - | Hotwords, format: word1 weight1 word2 weight2 |
response_format |
string | verbose_json |
Output format |
prompt |
string | - | Prompt text (reserved) |
temperature |
float | 0 |
Sampling temperature (reserved) |
Audio / Video Input Methods:
- File Upload: Use
fileparameter to upload an audio file or a video container with an audio track - URL / Local Path: Use
audio_addressparameter to provide an audio/video URL or server-local path, service will read it automatically - Precedence: If both
fileandaudio_addressare provided, the service usesfileand ignoresaudio_address
Usage Examples:
# Using OpenAI SDK
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="your_api_key")
with open("audio.wav", "rb") as f:
transcript = client.audio.transcriptions.create(
file=f,
response_format="verbose_json" # Get segments and speaker info
)
print(transcript.text)
# Using curl
curl -X POST "http://localhost:8000/v1/audio/transcriptions" \
-H "Authorization: Bearer your_api_key" \
-F "file=@audio.wav" \
-F "model=qwen3-asr-0.6b" \
-F "response_format=verbose_json" \
-F "enable_speaker_diarization=true" \
-F "enable_speaker_identification=true" \
-F "enable_text_cleanup=true" \
-F "hotwords=Qwen 2.0 ModelScope 1.5"
Supported Response Formats: json, text, srt, vtt, verbose_json
Alibaba Cloud Compatible API
| Endpoint | Method | Function |
|---|---|---|
/stream/v1/asr |
POST | Speech recognition (long audio support) |
/stream/v1/asr/models |
GET | Declared model/capability entries |
/stream/v1/asr/health |
GET | Health check |
/ws/v1/asr |
WebSocket | Qwen3-ASR streaming |
/ws/v1/asr/qwen |
WebSocket | Qwen3-ASR streaming (explicit path) |
/ws/v1/asr/funasr |
WebSocket | Removed; returns a deprecation error and asks clients to switch to /ws/v1/asr/qwen |
Request Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
audio_address |
string | https://media.cdn.vect.one/podcast_demo.mp4 (docs example) |
Audio/video URL, file://, or server-local path (optional; ignored when body content is uploaded) |
sample_rate |
int | 16000 |
Sample rate |
enable_speaker_diarization |
bool | true |
Enable speaker diarization |
enable_speaker_identification |
bool | true |
Match registered speaker database when diarization is enabled |
enable_text_cleanup |
bool | true |
Enable text deduplication, boundary-overlap trimming, and filler cleanup |
word_timestamps |
bool | false |
Return word-level timestamps when the backend supports them. Qwen CUDA vLLM and CPU Rust automatically use the forced aligner when enabled. |
vocabulary_id |
string | - | Hotwords (format: word1 weight1 word2 weight2) |
Usage Examples:
# Basic usage
curl -X POST "http://localhost:8000/stream/v1/asr" \
-H "Content-Type: application/octet-stream" \
--data-binary @audio.wav
# With parameters
curl -X POST "http://localhost:8000/stream/v1/asr?enable_speaker_diarization=true&enable_speaker_identification=true&enable_text_cleanup=true&vocabulary_id=Qwen%202.0%20ModelScope%201.5" \
-H "Content-Type: application/octet-stream" \
--data-binary @audio.wav
Meeting Offline API
| Endpoint | Method | Function |
|---|---|---|
/api/v1/asr/transcriptions |
POST | Create an offline meeting transcription task |
/api/v1/asr/transcriptions/{task_id} |
GET | Query task status and result |
audio_address is the required production input field for this endpoint.
{
"audio_address": "https://example.com/media/meeting.mp4",
"config": {
"enable_speaker": true,
"match_speaker_registry": true,
"enable_text_cleanup": true,
"speaker_threshold": 0.6,
"word_timestamps": false,
"hotwords": [
{ "hotword": "Qwen", "weight": 2.0 },
{ "hotword": "ModelScope", "weight": 1.5 }
]
}
}
Response Example:
{
"task_id": "xxx",
"status": 200,
"message": "SUCCESS",
"result": "Speaker1 content...\nSpeaker2 content...",
"duration": 60.5,
"processing_time": 1.234,
"segments": [
{
"text": "Today is a nice day.",
"start_time": 0.0,
"end_time": 2.5,
"speaker_id": "Speaker1",
"word_tokens": [
{"text": "Today", "start_time": 0.0, "end_time": 0.5},
{"text": "is", "start_time": 0.5, "end_time": 0.7},
{"text": "a nice day", "start_time": 0.7, "end_time": 1.5}
]
}
]
}
Speaker Diarization
Multi-speaker automatic identification based on CAM++ model:
- Enabled by Default -
enable_speaker_diarization=true - Automatic Detection - No preset speaker count needed, model auto-detects
- Speaker Labels - Response includes
speaker_idfield (e.g., "Speaker1", "Speaker2") - Smart Merging - Two-layer merge strategy to avoid isolated short segments:
- Layer 1: Accumulate merge same-speaker segments < 10 seconds
- Layer 2: Accumulate merge continuous segments up to 60 seconds
- Subtitle Support - SRT/VTT output includes speaker labels
[Speaker1] text content
Disable speaker diarization:
# OpenAI API
-F "enable_speaker_diarization=false"
# Alibaba Cloud API
?enable_speaker_diarization=false
Audio Processing
Intelligent Segmentation Strategy
Automatic long audio segmentation:
- VAD Voice Detection - Detect voice boundaries, filter silence
- Greedy Merge - Accumulate voice segments, ensure each segment does not exceed
MAX_SEGMENT_SEC(default 60s) - Silence Split - Force split when silence between voice segments exceeds 3 seconds
- Batch Inference - Multi-segment parallel processing, 2-3x performance improvement in GPU mode
WebSocket Streaming Limitations
Qwen3-ASR Streaming (using /ws/v1/asr or /ws/v1/asr/qwen):
- ✅ Multi-language real-time recognition
- ✅ CUDA vLLM and CPU Rust both support the current streaming path
- ❌ Word-level timestamps are not available in the current streaming path
Qwen3 Runtime Matrix
| Runtime | Backend | Offline | WebSocket Streaming | Word Timestamps Offline | Word Timestamps Streaming | Maturity |
|---|---|---|---|---|---|---|
| Linux + NVIDIA GPU | Official vLLM 0.20.0 | ✅ | ✅ | ✅ | ❌ | Production-oriented |
| CPU / macOS | QwenASR Rust | ✅ | ✅ | ✅ (forced aligner) | ❌ | Recommended local fallback |
Offline-Capable Models
| Model ID | Name | Description | Features |
|---|---|---|---|
qwen3-asr-1.7b |
Qwen3-ASR 1.7B | High-performance multilingual ASR, 52 languages + dialects; CUDA uses vLLM | Offline/Realtime |
qwen3-asr-0.6b |
Qwen3-ASR 0.6B | Lightweight multilingual ASR; CUDA uses vLLM, CPU/macOS uses Rust backend | Offline/Realtime |
Runtime selection:
- VRAM >= 32GB: Select
qwen3-asr-1.7b - VRAM < 32GB: Select
qwen3-asr-0.6b - No CUDA: Select the vendored Rust-backed
qwen3-asr-0.6b - macOS / Apple Silicon: Always default to
qwen3-asr-0.6b, regardless of memory size - Environment override: Set
QWEN3_ASR_MODEL=qwen3-asr-1.7borQWEN3_ASR_MODEL=qwen3-asr-0.6bto bypass automatic selection
At startup the service checks the current runtime model plan and downloads missing models from ModelScope by default.
Environment Variables
Recommended public settings:
| Variable | Default | Description |
|---|---|---|
API_KEY |
- | API authentication key (optional, unauthenticated if not set) |
LOG_LEVEL |
INFO |
Log level (DEBUG/INFO/WARNING/ERROR) |
MAX_AUDIO_SIZE |
2048 |
Max audio file size (MB, supports units like 2GB) |
ASR_BATCH_SIZE |
4 |
ASR batch size for long-audio segment processing |
MAX_SEGMENT_SEC |
60 |
Max audio segment duration (seconds) |
ASR_ENABLE_NEARFIELD_FILTER |
true |
Enable far-field sound filtering |
QWEN3_ASR_MODEL |
auto | Force qwen3-asr-1.7b or qwen3-asr-0.6b instead of VRAM-based selection |
QWEN_GPU_MEMORY_UTILIZATION |
0.9 |
Upper bound for vLLM GPU memory reservation; lower it on shared GPUs, raise it when KV cache is too small |
QWEN_VLLM_ENFORCE_EAGER |
true |
Force vLLM eager execution for compatibility; set false to allow CUDA Graph optimization on supported NVIDIA deployments |
Far-field filter notes:
ASR_NEARFIELD_RMS_THRESHOLD=0.01is the current default and recommended starting point- raise it in noisy rooms to filter more background speech
- lower it in quiet rooms if soft speech is being dropped
- use
LOG_LEVEL=DEBUGtemporarily when you need to inspect filter behavior
Advanced backend-specific settings:
| Variable | Default | Description |
|---|---|---|
QWEN_RUST_CPU_WORKERS |
4 |
CPU Rust backend worker count (Rust ASR / forced align default to 4 runtimes) |
QWENASR_LIBRARY_PATH |
auto-detect | Override vendored Rust dylib/so path |
Resource Requirements
Minimum (CPU):
- CPU: 4 cores
- Memory: 16GB
- Disk: 20GB
Recommended (GPU):
- CPU: 4 cores
- Memory: 16GB
- GPU: NVIDIA GPU (16GB+ VRAM)
- Disk: 20GB
API Documentation
After starting the service:
- Swagger UI:
http://localhost:8000/docs - ReDoc:
http://localhost:8000/redoc
Links
- Deployment Guide: Detailed Docs
- Qwen3-ASR: Qwen3-ASR GitHub
- FunASR: FunASR GitHub
- Chinese README: 中文文档
License
This project uses the MIT License - see LICENSE file for details.
Star History
Contributing
Issues and Pull Requests are welcome to improve the project!
