60 lines
2.1 KiB
Markdown
60 lines
2.1 KiB
Markdown
# FunASR realtime browser demo
|
|
|
|
The frontend and backend start separately. The backend launcher loads three
|
|
required local assets: streaming ASR, FSMN VAD, and CAM++ speaker verification.
|
|
It starts the CAM++ model service and the WebSocket service and stops both
|
|
together. No model is downloaded by the launcher. Use the shared model manifest
|
|
and downloader to prepare the three required snapshots:
|
|
|
|
~~~powershell
|
|
python scripts/download_models.py --funasr-runtime
|
|
~~~
|
|
|
|
## Model directories
|
|
|
|
Put the assets under models/ or set MODEL_DIR in .env. The default
|
|
FunASR names resolve to these local directories:
|
|
|
|
- models/iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online
|
|
- models/damo/speech_fsmn_vad_zh-cn-16k-common-pytorch
|
|
- models/iic/speech_campplus_sv_zh-cn_16k-common (or the damo/ variant)
|
|
|
|
If your directories have different names, set FUNASR_ASR_MODEL and
|
|
FUNASR_VAD_MODEL to their paths. The CAM++ directories follow
|
|
model_manifest.json; set CAM_MODEL_PATH for another location. Backend startup reports every checked path when a
|
|
required asset is missing.
|
|
|
|
## Start
|
|
|
|
First install a torch/torchaudio build suitable for the host CPU or CUDA,
|
|
then install the project dependencies:
|
|
|
|
~~~powershell
|
|
cd D:\github-project\ASR\Asr-demo
|
|
python -m pip install -r requirements-funasr.txt
|
|
python -m pip install -r requirements-auxiliary.txt
|
|
python scripts/download_models.py --funasr-runtime
|
|
if (-not (Test-Path .env)) { Copy-Item .env.funasr.example .env }
|
|
~~~
|
|
|
|
In one terminal start the backend:
|
|
|
|
~~~powershell
|
|
python scripts\run_backend.py
|
|
~~~
|
|
|
|
In another terminal start the frontend:
|
|
|
|
~~~powershell
|
|
python scripts\run_frontend.py
|
|
~~~
|
|
|
|
Open http://127.0.0.1:8080/. The backend WebSocket listens on port 8082 and
|
|
the CAM++ service on port 8010. BACKEND_PUBLIC_URL must be reachable from
|
|
the browser. Set FRONTEND_ORIGIN to the exact frontend origin if it differs
|
|
from the default.
|
|
|
|
The browser sends start, PCM16 frames, and stop/eof over WebSocket. Speaker
|
|
labels are required. Short or unusable speech may still receive an unknown
|
|
speaker label, but missing CAM++ prevents backend startup.
|