ASR-demo/FUNASR_README.md

60 lines
2.1 KiB
Markdown

# FunASR realtime browser demo
The frontend and backend start separately. The backend launcher loads three
required local assets: streaming ASR, FSMN VAD, and CAM++ speaker verification.
It starts the CAM++ model service and the WebSocket service and stops both
together. No model is downloaded by the launcher. Use the shared model manifest
and downloader to prepare the three required snapshots:
~~~powershell
python scripts/download_models.py --funasr-runtime
~~~
## Model directories
Put the assets under models/ or set MODEL_DIR in .env. The default
FunASR names resolve to these local directories:
- models/iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online
- models/damo/speech_fsmn_vad_zh-cn-16k-common-pytorch
- models/iic/speech_campplus_sv_zh-cn_16k-common (or the damo/ variant)
If your directories have different names, set FUNASR_ASR_MODEL and
FUNASR_VAD_MODEL to their paths. The CAM++ directories follow
model_manifest.json; set CAM_MODEL_PATH for another location. Backend startup reports every checked path when a
required asset is missing.
## Start
First install a torch/torchaudio build suitable for the host CPU or CUDA,
then install the project dependencies:
~~~powershell
cd D:\github-project\ASR\Asr-demo
python -m pip install -r requirements-funasr.txt
python -m pip install -r requirements-auxiliary.txt
python scripts/download_models.py --funasr-runtime
if (-not (Test-Path .env)) { Copy-Item .env.funasr.example .env }
~~~
In one terminal start the backend:
~~~powershell
python scripts\run_backend.py
~~~
In another terminal start the frontend:
~~~powershell
python scripts\run_frontend.py
~~~
Open http://127.0.0.1:8080/. The backend WebSocket listens on port 8082 and
the CAM++ service on port 8010. BACKEND_PUBLIC_URL must be reachable from
the browser. Set FRONTEND_ORIGIN to the exact frontend origin if it differs
from the default.
The browser sends start, PCM16 frames, and stop/eof over WebSocket. Speaker
labels are required. Short or unusable speech may still receive an unknown
speaker label, but missing CAM++ prevents backend startup.