44 lines
2.8 KiB
Markdown
44 lines
2.8 KiB
Markdown
# FunASR realtime browser demo
|
|
|
|
The frontend and backend start separately. The backend starts these supervised processes:
|
|
|
|
- FunASR native online WebSocket server, using the local streaming ASR and FSMN VAD models. The browser adapter sends FunASR's `mode=online`, chunk/look-back settings, fixed 60 ms PCM frames, and `is_speaking=false` end-of-input flush.
|
|
- CAM++ auxiliary service, which assigns stable speaker labels to finalized utterances.
|
|
- A small browser protocol adapter that translates the unchanged Tencent demo message format to FunASR's native WS format. It does not run a second ASR/VAD segmentation pipeline.
|
|
|
|
The frontend serves the exact files from the local `tencent-demo/static` directory and proxies the page's same-origin `/ws` and `/api/stop` requests to the backend.
|
|
|
|
## Model directories
|
|
|
|
Put assets under `models/` or set `MODEL_DIR` in `.env`. The default names resolve to these local directories:
|
|
|
|
- `models/iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online`
|
|
- `models/damo/speech_fsmn_vad_zh-cn-16k-common-pytorch`
|
|
- `models/iic/speech_campplus_sv_zh-cn_16k-common` (or the configured `damo/` variant)
|
|
|
|
If directories have different names, set `FUNASR_ASR_MODEL` and `FUNASR_VAD_MODEL` to their local paths. CAM++ follows `model_manifest.json`; set `CAM_MODEL_PATH` for another location. Startup checks all three assets before exposing the browser bridge.
|
|
|
|
## Install and start
|
|
|
|
Install a torch/torchaudio build suitable for the host, then install the project dependencies and models:
|
|
|
|
~~~powershell
|
|
python -m pip install -r backend/requirements-funasr.txt
|
|
python -m pip install -r backend/requirements-auxiliary.txt
|
|
python scripts/download_models.py --funasr-runtime
|
|
if (-not (Test-Path .env)) { Copy-Item .env.funasr.example .env }
|
|
~~~
|
|
|
|
Start the backend and frontend in separate terminals:
|
|
|
|
~~~powershell
|
|
python backend\run_backend.py
|
|
python frontend\run_frontend.py
|
|
~~~
|
|
|
|
Defaults are frontend port 8080, browser backend port 8082, CAM++ HTTP port 8010, and native FunASR WS port 10095 bound to loopback. Change `FRONTEND_PORT`, `WEB_PORT`, `AUXILIARY_SERVICE_URL`, `FUNASR_NATIVE_WS_HOST`, and `FUNASR_NATIVE_WS_PORT` together when needed. `BACKEND_INTERNAL_URL` is the backend origin reachable from the frontend process; it defaults to `http://127.0.0.1:${WEB_PORT}`.
|
|
|
|
`FUNASR_DEVICE` and `FUNASR_VAD_DEVICE` control ASR and VAD placement independently. Set `AUXILIARY_DEVICE=cpu` when CAM++ should not share the ASR GPU. `FUNASR_CHUNK_SIZE` and `FUNASR_CHUNK_INTERVAL` control FunASR's native chunk protocol; the default `[0,10,5]` and interval 10 send the current chunk in 600 ms groups.
|
|
|
|
Real-time file input currently accepts raw PCM16 or 16 kHz mono PCM WAV, matching the backend's available decoder. Speaker labels are computed for every finalized utterance; very short or silent segments can still be marked as unknown by CAM++.
|