ASR-demo/FUNASR_README.md

2.8 KiB

FunASR realtime browser demo

The frontend and backend start separately. The backend starts these supervised processes:

  • FunASR native online WebSocket server, using the local streaming ASR and FSMN VAD models. The browser adapter sends FunASR's mode=online, chunk/look-back settings, fixed 60 ms PCM frames, and is_speaking=false end-of-input flush.
  • CAM++ auxiliary service, which assigns stable speaker labels to finalized utterances.
  • A small browser protocol adapter that translates the unchanged Tencent demo message format to FunASR's native WS format. It does not run a second ASR/VAD segmentation pipeline.

The frontend serves the exact files from the local tencent-demo/static directory and proxies the page's same-origin /ws and /api/stop requests to the backend.

Model directories

Put assets under models/ or set MODEL_DIR in .env. The default names resolve to these local directories:

  • models/iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online
  • models/damo/speech_fsmn_vad_zh-cn-16k-common-pytorch
  • models/iic/speech_campplus_sv_zh-cn_16k-common (or the configured damo/ variant)

If directories have different names, set FUNASR_ASR_MODEL and FUNASR_VAD_MODEL to their local paths. CAM++ follows model_manifest.json; set CAM_MODEL_PATH for another location. Startup checks all three assets before exposing the browser bridge.

Install and start

The root requirements.txt includes FunASR and auxiliary-service dependencies plus the Qwen vLLM stack pinned for a GB10 CUDA 13 host. Adjust those CUDA-specific pins for other hosts before installing:

python -m pip install -r requirements.txt
python scripts/download_models.py --funasr-runtime
if (-not (Test-Path .env)) { Copy-Item .env.funasr.example .env }

Start the backend and frontend in separate terminals:

python backend\run_backend.py
python frontend\run_frontend.py

Defaults are frontend port 8080, browser backend port 8082, CAM++ HTTP port 8010, and native FunASR WS port 10095 bound to loopback. Change FRONTEND_PORT, WEB_PORT, AUXILIARY_SERVICE_URL, FUNASR_NATIVE_WS_HOST, and FUNASR_NATIVE_WS_PORT together when needed. BACKEND_INTERNAL_URL is the backend origin reachable from the frontend process; it defaults to http://127.0.0.1:${WEB_PORT}.

FUNASR_DEVICE and FUNASR_VAD_DEVICE control ASR and VAD placement independently. Set AUXILIARY_DEVICE=cpu when CAM++ should not share the ASR GPU. FUNASR_CHUNK_SIZE and FUNASR_CHUNK_INTERVAL control FunASR's native chunk protocol; the default [0,10,5] and interval 10 send the current chunk in 600 ms groups.

Real-time file input currently accepts raw PCM16 or 16 kHz mono PCM WAV, matching the backend's available decoder. Speaker labels are computed for every finalized utterance; very short or silent segments can still be marked as unknown by CAM++.