# FunASR realtime browser demo The frontend and backend start separately. The backend starts these supervised processes: - FunASR native online WebSocket server, using the local streaming ASR and FSMN VAD models. The browser adapter sends FunASR's `mode=online`, chunk/look-back settings, fixed 60 ms PCM frames, and `is_speaking=false` end-of-input flush. - CAM++ auxiliary service, which assigns stable speaker labels to finalized utterances. - A small browser protocol adapter that translates the unchanged Tencent demo message format to FunASR's native WS format. It does not run a second ASR/VAD segmentation pipeline. The frontend serves the exact files from the local `tencent-demo/static` directory and proxies the page's same-origin `/ws` and `/api/stop` requests to the backend. ## Model directories Put assets under `models/` or set `MODEL_DIR` in `.env`. The default names resolve to these local directories: - `models/iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online` - `models/damo/speech_fsmn_vad_zh-cn-16k-common-pytorch` - `models/iic/speech_campplus_sv_zh-cn_16k-common` (or the configured `damo/` variant) If directories have different names, set `FUNASR_ASR_MODEL` and `FUNASR_VAD_MODEL` to their local paths. CAM++ follows `model_manifest.json`; set `CAM_MODEL_PATH` for another location. Startup checks all three assets before exposing the browser bridge. ## Install and start Install a torch/torchaudio build suitable for the host, then install the project dependencies and models: ~~~powershell python -m pip install -r backend/requirements-funasr.txt python -m pip install -r backend/requirements-auxiliary.txt python scripts/download_models.py --funasr-runtime if (-not (Test-Path .env)) { Copy-Item .env.funasr.example .env } ~~~ Start the backend and frontend in separate terminals: ~~~powershell python backend\run_backend.py python frontend\run_frontend.py ~~~ Defaults are frontend port 8080, browser backend port 8082, CAM++ HTTP port 8010, and native FunASR WS port 10095 bound to loopback. Change `FRONTEND_PORT`, `WEB_PORT`, `AUXILIARY_SERVICE_URL`, `FUNASR_NATIVE_WS_HOST`, and `FUNASR_NATIVE_WS_PORT` together when needed. `BACKEND_INTERNAL_URL` is the backend origin reachable from the frontend process; it defaults to `http://127.0.0.1:${WEB_PORT}`. `FUNASR_DEVICE` and `FUNASR_VAD_DEVICE` control ASR and VAD placement independently. Set `AUXILIARY_DEVICE=cpu` when CAM++ should not share the ASR GPU. `FUNASR_CHUNK_SIZE` and `FUNASR_CHUNK_INTERVAL` control FunASR's native chunk protocol; the default `[0,10,5]` and interval 10 send the current chunk in 600 ms groups. Real-time file input currently accepts raw PCM16 or 16 kHz mono PCM WAV, matching the backend's available decoder. Speaker labels are computed for every finalized utterance; very short or silent segments can still be marked as unknown by CAM++.