26 KiB
Qwen3-ASR
Ready-to-use Local Speech Recognition API Service
Speech recognition API service centered on Qwen3-ASR, with NVIDIA CUDA vLLM, MetaX/MuXi MACA vLLM, and CPU Rust backends, OpenAI API compatibility, Alibaba Cloud Speech API compatibility, and a Paraformer realtime websocket capability.
Live Demo Site
- Web Demo: https://asr.vect.one
Demo
Release 1.0.1
v1.0.1is the current patch release.v1.0.0introduced a large breaking refactor relative to the earliermainbranch. If you are upgrading frommain, read the release notes before reusing old deployment assumptions.Key breaking changes:
- Python dependency management is now
uv-based (pyproject.toml+uv.lock);requirements*.txtare gone- Runtime stack changed to
NVIDIA/MetaX GPU -> vLLM,CPU/macOS -> vendored QwenASR RustMLX/ Apple Silicon GPU path has been removed;mpsis normalized tocpu- macOS / Apple Silicon now defaults to
qwen3-asr-0.6b; setQWEN3_ASR_MODELto override itENABLED_MODELShas been removed
Features
- Hybrid Runtime Stack - Uses auto-selected Qwen3-ASR for offline inference and Paraformer realtime for websocket streaming
- Speaker Diarization - Automatic multi-speaker identification using CAM++ model
- OpenAI API Compatible - Supports
/v1/audio/transcriptionsendpoint, works with OpenAI SDK - Alibaba Cloud API Compatible - Supports Alibaba Cloud Speech RESTful API and WebSocket streaming protocol
- WebSocket Streaming - Real-time streaming speech recognition with low latency
- Smart Far-Field Filtering - Automatically filters far-field sounds and ambient noise in streaming ASR
- Intelligent Audio Segmentation - VAD-based greedy merge algorithm for automatic long audio splitting
- GPU Batch Processing - Batch inference support, 2-3x faster than sequential processing
- Resource-Aware Runtime - Auto-selects the appropriate Qwen3-ASR model for the current machine
Acknowledgements
- Qwen3-ASR provides the official model family and multimodal/vLLM usage guidance
- QwenASR provides the CPU Rust backend vendored by this project
Quick Deployment
1. Docker Deployment (Recommended)
# Copy and edit configuration
cp .env.example .env
# Edit .env to set API_KEY (optional)
# Compose defaults:
# /opt/dep/asr/models -> /app/models
# /opt/dep/asr/data -> /app/data
# /opt/dep/asr/data/logs, temp, tasks live under this data mount
# Optional: override any host mount root in .env
# export MODEL_STORAGE_DIR=/data/qwen3-asr-models
# export DATA_STORAGE_DIR=/data/qwen3-asr-data
# Start service (NVIDIA GPU version)
docker-compose up -d
# Or MetaX/MuXi GPU version
docker-compose -f docker-compose-metax.yml up -d
# Or Iluvatar/Tianshu GPU version
docker-compose -f docker-compose-iluvatar.yml up -d
# Or Moore Threads / MUSA GPU version
docker-compose -f docker-compose-mthreads.yml up -d
# Or CPU version
docker-compose -f docker-compose-cpu.yml up -d
# NVIDIA multi-GPU auto mode (one instance per visible GPU)
CUDA_VISIBLE_DEVICES=0,1,2,3 docker-compose up -d
# MetaX/MuXi multi-GPU auto mode
METAX_VISIBLE_DEVICES=0,1 docker-compose -f docker-compose-metax.yml up -d
# Iluvatar/Tianshu multi-GPU auto mode
ILUVATAR_VISIBLE_DEVICES=0,1 docker-compose -f docker-compose-iluvatar.yml up -d
# Moore Threads / MUSA multi-GPU auto mode
MTHREADS_VISIBLE_DEVICES=0,1 docker-compose -f docker-compose-mthreads.yml up -d
Service URLs:
- API Endpoint:
http://localhost:17003 - API Docs:
http://localhost:17003/docs
Optional built-in rate limit settings:
NGINX_RATE_LIMIT_RPS(global requests/sec,0= disabled)NGINX_RATE_LIMIT_BURST(global burst,0= auto use RPS)
docker run (alternative):
# NVIDIA GPU version
docker run -d --name qwen3-asr \
--gpus all \
-p 17003:8000 \
-e ACCELERATOR=nvidia \
-e CUDA_VISIBLE_DEVICES=0,1,2,3 \
-e API_KEY=your_api_key \
-v /opt/dep/asr/models:/app/models \
-v /opt/dep/asr/data:/app/data \
unis/qwen3-asr:gpu-latest
# MetaX/MuXi GPU version
docker run -d --name qwen3-asr-metax \
--privileged \
--network=host \
--pid=host \
--ipc=host \
-v /dev:/dev \
-v /opt/mxdriver:/opt/mxdriver:ro \
-e ACCELERATOR=metax \
-e PORT=17003 \
-e METAX_VISIBLE_DEVICES=0 \
-v /opt/dep/asr/models:/app/models \
-v /opt/dep/asr/data:/app/data \
unis/qwen3-asr:metax-latest
# CPU version
docker run -d --name qwen3-asr \
-p 17003:8000 \
-v /opt/dep/asr/models:/app/models \
-v /opt/dep/asr/data:/app/data \
unis/qwen3-asr:cpu-latest
Note: NVIDIA GPU images default to CUDA 13.0/cu130 with
torch 2.11.0+vllm 0.20.0. Developers can rebuildDockerfile.gpufor CUDA 12.6, CUDA 13.0, or another backend by overriding Docker build args. MetaX/MuXi images useDockerfile.metaxon top of an official MetaX vLLM image. In field deployments, use host networking plus privileged/devand/opt/mxdrivermounts so bothmx-smiand the MetaX PyTorch runtime can initialize devices. CPU images now supportqwen3-asr-0.6bvia the bundled QwenASR Rust backend. The default CPU image uses a portable Rust target; setQWENASR_RUST_TARGET_CPU=nativeonly for self-built, host-specific images. On CUDA vLLM and CPU Rust,word_timestamps=truenow triggers the forced aligner automatically. On macOS / Apple Silicon, Qwen3-ASR now runs through the Rust CPU backend.start.pynow forces the vLLM multiprocessing method tospawnso startup does not hit CUDA re-initialization failures in forked subprocesses.
Custom GPU backend builds:
# Default GPU build: CUDA 13.0 / PyTorch cu130
docker build -t qwen3-asr:gpu-cu130 -f Dockerfile.gpu .
# CUDA 12.6 build for older deployments
docker build -t qwen3-asr:gpu-cu126 -f Dockerfile.gpu \
--build-arg PYTORCH_BASE_IMAGE=pytorch/pytorch:2.11.0-cuda12.6-cudnn9-runtime \
--build-arg PYTORCH_CUDA_INDEX=https://download.pytorch.org/whl/cu126 \
--build-arg CUDA_NVCC_PACKAGE=cuda-nvcc-12-6 \
--build-arg TORCH_CUDA_ARCH_LIST="8.0;8.6;8.9" \
.
# CUDA 13.0 build when your driver/toolchain requires it
docker build -t qwen3-asr:gpu-cu130 -f Dockerfile.gpu \
--build-arg PYTORCH_BASE_IMAGE=pytorch/pytorch:2.11.0-cuda13.0-cudnn9-runtime \
--build-arg PYTORCH_CUDA_INDEX=https://download.pytorch.org/whl/cu130 \
--build-arg CUDA_NVCC_PACKAGE=cuda-nvcc-13-0 \
--build-arg TORCH_CUDA_ARCH_LIST="12.0+PTX" \
.
# MetaX/MuXi build: fuse this project into an official MetaX vLLM image
./scripts/package_vendor_gpu_image.sh \
--vendor metax \
--base-image <official-metax-vllm-image> \
-v n260-3.7.0.38
# Iluvatar/Tianshu build: fuse this project into the official Iluvatar vLLM image
docker pull registry.iluvatar.com.cn:10443/customer/sz/vllm0.17.0-4.4.0-x86:v5
./scripts/package_vendor_gpu_image.sh \
--vendor iluvatar \
--base-image registry.iluvatar.com.cn:10443/customer/sz/vllm0.17.0-4.4.0-x86:v5 \
-v vllm0.17.0-4.4.0-v5
# Moore Threads / MUSA build: fuse this project into the official MUSA vLLM image
docker pull registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519
./scripts/package_vendor_gpu_image.sh \
--vendor mthreads \
--base-image registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519 \
-v s4000_4.3.5_d0519
For MetaX/MuXi offline delivery, see docs/metax_offline_deployment.md. For Iluvatar/Tianshu offline delivery, see docs/iluvatar_offline_deployment.md. For Moore Threads / MUSA offline delivery, see docs/mthreads_offline_deployment.md.
Offline Deployment: You can now build a timestamped offline delivery folder that includes the image archive, compose file, env template, host-dir init script, and usage docs. The export script uses plain docker build + docker save, so it does not depend on buildx:
# 1. Build an offline delivery folder
./export_offline_bundle.sh --type gpu
# or
./export_offline_bundle.sh --type cpu
# or MetaX/MuXi GPU
./export_offline_bundle.sh \
--type metax \
--metax-base cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.17.0-maca.ai3.5.3.307-torch2.8-py312-ubuntu22.04-amd64 \
--skip-models
# or Iluvatar/Tianshu GPU
./export_offline_bundle.sh --type iluvatar --iluvatar-base registry.iluvatar.com.cn:10443/customer/sz/vllm0.17.0-4.4.0-x86:v5
# or Moore Threads / MUSA GPU
./export_offline_bundle.sh --type mthreads --mthreads-base registry.mthreads.com/presale/devtech/vllm_musa:s4000_4.3.5_d0519
# or build both in one bundle
./export_offline_bundle.sh --type all
# 2. Prepare models separately, without deleting existing model files
./scripts/download-models.sh --models-dir /opt/dep/asr/models
# 3. Copy the generated folder to the offline server
scp -r build-file/<timestamp>-all user@server:/opt/dep/asr/
# 4. On the offline server
cd /opt/dep/asr/<timestamp>-all
./init_host_dirs.sh
gunzip -c qwen3-asr-gpu-<timestamp>-amd64.tar.gz | docker load
gunzip -c qwen3-asr-cpu-<timestamp>-amd64.tar.gz | docker load
# NVIDIA GPU
docker compose up -d
# or MetaX/MuXi GPU
# docker compose -f docker-compose-metax.yml up -d
# or Iluvatar/Tianshu GPU
# docker compose -f docker-compose-iluvatar.yml up -d
# or Moore Threads / MUSA GPU
# docker compose -f docker-compose-mthreads.yml up -d
# or CPU
# docker compose -f docker-compose-cpu.yml up -d
Detailed deployment instructions: Deployment Guide
2. Non-Docker Deployment (Linux Server)
Run the FastAPI service directly on the server when you want to test without Docker. The commands below use NVIDIA as an example; use the matching scripts/sync_*_env.sh script for CPU or a vendor GPU backend (see the table in Local Development).
Requirements: Python 3.10–3.12, uv, FFmpeg, and the server's supported accelerator driver/runtime. The default NVIDIA environment installs the CUDA 13.0 PyTorch/vLLM stack from the project lock file.
# From the project root
python3 --version
uv --version
# Install the NVIDIA environment. For CPU, use ./scripts/sync_cpu_env.sh.
./scripts/sync_gpu_env.sh
# Copy the sample settings, then edit .env if needed.
cp .env.example .env
Set HOST=0.0.0.0 and PORT=8000 in .env if you want to access the service from another machine. Set API_KEY to enable authentication; it is optional for a local test. Model files use ./models by default. To store them elsewhere, set MODELS_DIR, MODELSCOPE_CACHE, and MODELSCOPE_PATH in .env as described in the sample file.
Download the models before starting so the first server launch does not wait for model downloads:
./scripts/download-models.sh --models-dir ./models --python-bin ./.venv/bin/python --mode local
This step requires access to ModelScope. If you skip it, start.py checks for missing models and attempts to download them during startup.
Start the service in the foreground and watch the startup logs:
./.venv/bin/python start.py
The default endpoint is http://<server-ip>:8000; open port 8000 in the server firewall if you are connecting remotely. Check the service and open the interactive API docs at http://<server-ip>:8000/docs. To test transcription with an audio file:
curl -X POST http://127.0.0.1:8000/v1/audio/transcriptions \
-H "Authorization: Bearer your_api_key" \
-F "file=@/path/to/test.wav"
Replace your_api_key with the value in .env. If API_KEY is unset, omit the Authorization header. Stop the foreground process with Ctrl+C.
Local Development
System Requirements:
- Python 3.10+
- CUDA 13.0+ for the default GPU image; CUDA 12.6 / 13.0 can be built with Docker args
- FFmpeg (audio format conversion)
Installation:
Runtime dependency locks now default to the GPU stack at the repo root, with CPU kept as a specialized environment:
| Mode | Command | Notes |
|---|---|---|
| NVIDIA GPU (default) | uv sync or ./scripts/sync_gpu_env.sh |
Syncs the root pyproject.toml and uv.lock into .venv, including CUDA 13.0/cu130 torch 2.11.0 / torchaudio 2.11.0 / torchvision 0.26.0 / vllm 0.20.0 |
| MetaX/MuXi GPU | ./scripts/sync_metax_env.sh |
Syncs common dependencies from environments/metax/pyproject.toml; optional GPU-stack install uses the MetaX MACA PyPI index with --no-deps by default |
| Iluvatar/Tianshu GPU | ./scripts/sync_iluvatar_env.sh |
Syncs common dependencies from environments/iluvatar/pyproject.toml; GPU stack should come from the official Iluvatar vLLM image |
| Moore Threads / MUSA GPU | ./scripts/sync_mthreads_env.sh |
Syncs common dependencies from environments/mthreads/pyproject.toml; GPU stack should come from the official Moore Threads MUSA vLLM image |
| CPU (specialized) | ./scripts/sync_cpu_env.sh |
Syncs the dedicated CPU lock in environments/cpu/pyproject.toml into .venv |
| Auto | ./scripts/sync_accel_env.sh |
Chooses MetaX when mx-smi is present, Iluvatar when ixsmi is present, Moore Threads when mthreads-gmi is present, otherwise NVIDIA when nvidia-smi is present, otherwise CPU |
# Clone project
cd qwen3-asr
# Install dependencies (Linux/NVIDIA CUDA)
uv sync
# Start service
source .venv/bin/activate
python start.py
MetaX/MuXi local development:
./scripts/sync_metax_env.sh
source .venv/bin/activate
ACCELERATOR=metax python start.py
Local model storage defaults to ./models under the project root:
./models/
Qwen/
iic/
damo/
Override it when needed:
export MODELS_DIR=/data/qwen3-asr-models
export MODELSCOPE_CACHE=/data
export MODELSCOPE_PATH=$MODELS_DIR
macOS / Apple Silicon local development:
./scripts/sync_cpu_env.sh
source .venv/bin/activate
python start.py
Runtime Defaults
Current runtime behavior on the mainline codebase:
ACCELERATOR=autoresolves tometaxwhenmx-smireports devices, theniluvatarwhenixsmireports devices, thenmthreadswhenmthreads-gmireports devices, otherwisenvidiawhen NVIDIA CUDA is available, otherwisecpuDEVICE=autoresolves to the active accelerator device (cuda:0for NVIDIA/MetaX/Iluvatar GPU, otherwisecpu)DEVICE=mpsis normalized tocpuLinux + NVIDIA CUDAuses officialvLLMLinux + MetaX/MuXi MACAuses the MetaX-compatible PyTorch/vLLM stackLinux + Iluvatar/Tianshuuses the Iluvatar official vLLM image stackLinux + CPUuses vendoredQwenASRRustmacOS / Apple Siliconalso uses vendoredQwenASRRust- macOS / Apple Silicon defaults to
qwen3-asr-0.6b qwen3-asr-1.7bon macOS is only used whenQWEN3_ASR_MODEL=qwen3-asr-1.7bword_timestamps=trueworks on the current offline CUDA and CPU Rust paths- WebSocket streaming does not currently return word-level timestamps
- CAM++ speaker diarization remains required and still follows
DEVICE; on CPU its main hotspot is speaker verification embedding
API Endpoints
OpenAI Compatible API
| Endpoint | Method | Function |
|---|---|---|
/v1/audio/transcriptions |
POST | Audio transcription (OpenAI compatible) |
/v1/models |
GET | Offline model list |
Request Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
file |
file | Preferred when provided | Audio/video file |
audio_address |
string | Optional | Audio/video URL (HTTP/HTTPS), file://, or server-local path. Ignored when file is also provided |
language |
string | Auto-detect | Language code (zh/en/ja) |
enable_speaker_diarization |
bool | true |
Enable speaker diarization |
enable_speaker_identification |
bool | true |
Match registered speaker database when diarization is enabled |
enable_text_cleanup |
bool | true |
Enable text deduplication, boundary-overlap trimming, and filler cleanup |
word_timestamps |
bool | false |
Return word-level timestamps when the backend supports them. Qwen CUDA vLLM and CPU Rust automatically use the forced aligner when enabled. |
hotwords |
string | - | Hotwords, format: word1 weight1 word2 weight2 |
response_format |
string | verbose_json |
Output format |
prompt |
string | - | Prompt text (reserved) |
temperature |
float | 0 |
Sampling temperature (reserved) |
Audio / Video Input Methods:
- File Upload: Use
fileparameter to upload an audio file or a video container with an audio track - URL / Local Path: Use
audio_addressparameter to provide an audio/video URL or server-local path, service will read it automatically - Precedence: If both
fileandaudio_addressare provided, the service usesfileand ignoresaudio_address
Usage Examples:
# Using OpenAI SDK
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="your_api_key")
with open("audio.wav", "rb") as f:
transcript = client.audio.transcriptions.create(
file=f,
response_format="verbose_json" # Get segments and speaker info
)
print(transcript.text)
# Using curl
curl -X POST "http://localhost:8000/v1/audio/transcriptions" \
-H "Authorization: Bearer your_api_key" \
-F "file=@audio.wav" \
-F "model=qwen3-asr-0.6b" \
-F "response_format=verbose_json" \
-F "enable_speaker_diarization=true" \
-F "enable_speaker_identification=true" \
-F "enable_text_cleanup=true" \
-F "hotwords=Qwen 2.0 ModelScope 1.5"
Supported Response Formats: json, text, srt, vtt, verbose_json
Alibaba Cloud Compatible API
| Endpoint | Method | Function |
|---|---|---|
/stream/v1/asr |
POST | Speech recognition (long audio support) |
/stream/v1/asr/models |
GET | Declared model/capability entries |
/stream/v1/asr/health |
GET | Health check |
/ws/v1/asr |
WebSocket | Qwen3-ASR streaming |
/ws/v1/asr/qwen |
WebSocket | Qwen3-ASR streaming (explicit path) |
/ws/v1/asr/funasr |
WebSocket | Removed; returns a deprecation error and asks clients to switch to /ws/v1/asr/qwen |
Request Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
audio_address |
string | https://media.cdn.vect.one/podcast_demo.mp4 (docs example) |
Audio/video URL, file://, or server-local path (optional; ignored when body content is uploaded) |
sample_rate |
int | 16000 |
Sample rate |
enable_speaker_diarization |
bool | true |
Enable speaker diarization |
enable_speaker_identification |
bool | true |
Match registered speaker database when diarization is enabled |
enable_text_cleanup |
bool | true |
Enable text deduplication, boundary-overlap trimming, and filler cleanup |
word_timestamps |
bool | false |
Return word-level timestamps when the backend supports them. Qwen CUDA vLLM and CPU Rust automatically use the forced aligner when enabled. |
vocabulary_id |
string | - | Hotwords (format: word1 weight1 word2 weight2) |
Usage Examples:
# Basic usage
curl -X POST "http://localhost:8000/stream/v1/asr" \
-H "Content-Type: application/octet-stream" \
--data-binary @audio.wav
# With parameters
curl -X POST "http://localhost:8000/stream/v1/asr?enable_speaker_diarization=true&enable_speaker_identification=true&enable_text_cleanup=true&vocabulary_id=Qwen%202.0%20ModelScope%201.5" \
-H "Content-Type: application/octet-stream" \
--data-binary @audio.wav
Meeting Offline API
| Endpoint | Method | Function |
|---|---|---|
/api/v1/asr/transcriptions |
POST | Create an offline meeting transcription task |
/api/v1/asr/transcriptions/{task_id} |
GET | Query task status and result |
audio_address is the required production input field for this endpoint.
{
"audio_address": "https://example.com/media/meeting.mp4",
"config": {
"enable_speaker": true,
"match_speaker_registry": true,
"enable_text_cleanup": true,
"speaker_threshold": 0.6,
"word_timestamps": false,
"hotwords": [
{ "hotword": "Qwen", "weight": 2.0 },
{ "hotword": "ModelScope", "weight": 1.5 }
]
}
}
Response Example:
{
"task_id": "xxx",
"status": 200,
"message": "SUCCESS",
"result": "Speaker1 content...\nSpeaker2 content...",
"duration": 60.5,
"processing_time": 1.234,
"segments": [
{
"text": "Today is a nice day.",
"start_time": 0.0,
"end_time": 2.5,
"speaker_id": "Speaker1",
"word_tokens": [
{"text": "Today", "start_time": 0.0, "end_time": 0.5},
{"text": "is", "start_time": 0.5, "end_time": 0.7},
{"text": "a nice day", "start_time": 0.7, "end_time": 1.5}
]
}
]
}
Speaker Diarization
Multi-speaker automatic identification based on CAM++ model:
- Enabled by Default -
enable_speaker_diarization=true - Automatic Detection - No preset speaker count needed, model auto-detects
- Speaker Labels - Response includes
speaker_idfield (e.g., "Speaker1", "Speaker2") - Smart Merging - Two-layer merge strategy to avoid isolated short segments:
- Layer 1: Accumulate merge same-speaker segments < 10 seconds
- Layer 2: Accumulate merge continuous segments up to 60 seconds
- Subtitle Support - SRT/VTT output includes speaker labels
[Speaker1] text content
Disable speaker diarization:
# OpenAI API
-F "enable_speaker_diarization=false"
# Alibaba Cloud API
?enable_speaker_diarization=false
Audio Processing
Intelligent Segmentation Strategy
Automatic long audio segmentation:
- VAD Voice Detection - Detect voice boundaries, filter silence
- Greedy Merge - Accumulate voice segments, ensure each segment does not exceed
MAX_SEGMENT_SEC(default 60s) - Silence Split - Force split when silence between voice segments exceeds 3 seconds
- Batch Inference - Multi-segment parallel processing, 2-3x performance improvement in GPU mode
WebSocket Streaming Limitations
Qwen3-ASR Streaming (using /ws/v1/asr or /ws/v1/asr/qwen):
- ✅ Multi-language real-time recognition
- ✅ CUDA vLLM and CPU Rust both support the current streaming path
- ❌ Word-level timestamps are not available in the current streaming path
Qwen3 Runtime Matrix
| Runtime | Backend | Offline | WebSocket Streaming | Word Timestamps Offline | Word Timestamps Streaming | Maturity |
|---|---|---|---|---|---|---|
| Linux + NVIDIA GPU | Official vLLM 0.20.0 | ✅ | ✅ | ✅ | ❌ | Production-oriented |
| CPU / macOS | QwenASR Rust | ✅ | ✅ | ✅ (forced aligner) | ❌ | Recommended local fallback |
Offline-Capable Models
| Model ID | Name | Description | Features |
|---|---|---|---|
qwen3-asr-1.7b |
Qwen3-ASR 1.7B | High-performance multilingual ASR, 52 languages + dialects; CUDA uses vLLM | Offline/Realtime |
qwen3-asr-0.6b |
Qwen3-ASR 0.6B | Lightweight multilingual ASR; CUDA uses vLLM, CPU/macOS uses Rust backend | Offline/Realtime |
Runtime selection:
- Default: Use
qwen3-asr-0.6bfor both offline and realtime ASR. - No CUDA: Select the vendored Rust-backed
qwen3-asr-0.6b - macOS / Apple Silicon: Always default to
qwen3-asr-0.6b, regardless of memory size - Environment override: Set
QWEN3_ASR_MODEL=qwen3-asr-1.7bto use 1.7B instead of the shared 0.6B default.
At startup the service checks the current runtime model plan and downloads missing models from ModelScope by default.
Environment Variables
Recommended public settings:
| Variable | Default | Description |
|---|---|---|
API_KEY |
- | API authentication key (optional, unauthenticated if not set) |
LOG_LEVEL |
INFO |
Log level (DEBUG/INFO/WARNING/ERROR) |
MAX_AUDIO_SIZE |
2048 |
Max audio file size (MB, supports units like 2GB) |
ASR_BATCH_SIZE |
4 |
ASR batch size for long-audio segment processing |
MAX_SEGMENT_SEC |
60 |
Max audio segment duration (seconds) |
ASR_ENABLE_NEARFIELD_FILTER |
true |
Enable far-field sound filtering |
QWEN3_ASR_MODEL |
qwen3-asr-0.6b |
Select the shared offline/realtime model; set 1.7B to override |
QWEN_GPU_MEMORY_UTILIZATION |
0.9 |
Upper bound for vLLM GPU memory reservation; lower it on shared GPUs, raise it when KV cache is too small |
QWEN_VLLM_ENFORCE_EAGER |
true |
Force vLLM eager execution for compatibility; set false to allow CUDA Graph optimization on supported NVIDIA deployments |
Far-field filter notes:
ASR_NEARFIELD_RMS_THRESHOLD=0.01is the current default and recommended starting point- raise it in noisy rooms to filter more background speech
- lower it in quiet rooms if soft speech is being dropped
- use
LOG_LEVEL=DEBUGtemporarily when you need to inspect filter behavior
Advanced backend-specific settings:
| Variable | Default | Description |
|---|---|---|
QWEN_RUST_CPU_WORKERS |
4 |
CPU Rust backend worker count (Rust ASR / forced align default to 4 runtimes) |
QWENASR_LIBRARY_PATH |
auto-detect | Override vendored Rust dylib/so path |
Resource Requirements
Minimum (CPU):
- CPU: 4 cores
- Memory: 16GB
- Disk: 20GB
Recommended (GPU):
- CPU: 4 cores
- Memory: 16GB
- GPU: NVIDIA GPU (16GB+ VRAM)
- Disk: 20GB
API Documentation
After starting the service:
- Swagger UI:
http://localhost:8000/docs - ReDoc:
http://localhost:8000/redoc
Links
- Deployment Guide: Detailed Docs
- Qwen3-ASR: Qwen3-ASR GitHub
- FunASR: FunASR GitHub
- Chinese README: 中文文档
License
This project uses the MIT License - see LICENSE file for details.
Star History
Contributing
Issues and Pull Requests are welcome to improve the project!
