- C1: stop feeding the unconsumed barge-in TTSQueue in the reply path (deadlocked after 64 chunks on a real long reply) - C2: encode reply Opus at the TTS/codec sample rate (24 kHz), not hardcoded 16 kHz; encoder now rate-generic - M3: bound the mic PCM buffer to the 60 s max utterance - M4: tts_sentence_start per speakable sentence, not per LLM token - Minor: send tts.stop on a failed turn; constant-time token compare - Sec: untrack + gitignore config.yaml; add config.example.yaml with placeholder secrets - Tests: 23 pass (2 new regressions: long-reply deadlock, rate-generic encoder)
7.3 KiB
Plan — hermes-xiaozhi-bridge
Status: Phase 1.5 — live-verified (CPU host): 18 tests + live WS smoke PASS Last update: 2026-10-03 Next action: voice-server — wire real STT/TTS/Opus backends, verify protocol constants against firmware (CON-004).
Scope of this host's run
This Mac has no GPU / device / native libs, so the deliverable here is the testable bridge skeleton (AC-001..AC-005). GPU/STT/TTS/device behavior is stubbed behind interfaces and verified with mocks; the real backends are lazy imports that activate only on the voice-server.
Task breakdown
- TASK-001 Repo, venv, deps (fastapi/websockets/pydantic/pytest…) — DEC-001
- TASK-002
config.py+config.yaml(engines, devices, protocol constants) - TASK-003
xiaozhi/— protocol (hello/hello_reply frames), messages, session, websocket endpoint - TASK-004
audio/— buffer (ring), opus (lazy), VAD (energy + mock) end-of-speech - TASK-005
stt/— base + typhoon + whisper (lazy) + mock - TASK-006
tts/— base + jaitts (lazy) + chunker (Thai) + queue (barge-in cancel) + mock - TASK-007
hermes/— client, session (multi-turn, bounded ctx), voice_profile (reasoning none) - TASK-008
security/auth.py— device allowlist, token hash - TASK-009
gpu/— monitor, voice_models, resource_manager (voice-full/lite/off) - TASK-010
app/main.py— FastAPI app,/health,/ws/xiaozhi - TASK-011
metrics/latency.py— timestamps + End-of-Speech→First Audible Audio - TASK-012
tests/— suite green (AC-001..AC-005) - TASK-013 README, requirements.txt,
__init__.pyfiles, .gitignore (.env.example= N/A: config.yaml is the single source of truth)
Requirement coverage
| Req | Task(s) |
|---|---|
| REQ-001 | TASK-003, 010 |
| REQ-002 | TASK-004 |
| REQ-003 | TASK-005 |
| REQ-004 | TASK-007 (voice_profile) |
| REQ-005 | TASK-007 (session) |
| REQ-006 | TASK-006 |
| REQ-007 | TASK-006 (queue barge-in) |
| REQ-008 | TASK-007 (hermes client → session_search) |
| REQ-009 | TASK-008 |
| REQ-010 | TASK-009 |
| REQ-011 | TASK-011 |
| REQ-012 | TASK-002/008 (bind LAN; no CF logic) |
Verification
pytest -q(CPU-only, mock backends) → expect all green.python -c "import app.main"import check (AC-002).- Import-safety check: no GPU/opus/STT/TTS import at module load (CON-002).
Evidence
- 2026-10-03
pytest -q→ 18 passed (protocol, config, auth, chunker, barge-in queue, latency, voice-profile guard, GPU planner, VAD, e2e mock turn, app import + live WS handshake+turn). - 2026-10-03
python -c "import app.main"→ OK;sys.modulescheck → no torch/typhoon/whisper/jait/opus/numpy loaded (CON-002). - 2026-10-03 Live boot on Windows CPU host: fixed
main()ignoring config.yaml (wasConfig()defaults); port 8765 taken by another local service → 8766. Issued real token forxiaozhi-main(hash in config.yaml).smoke_live.pyagainst running server: health ok, bad-token rejected, bad-device rejected, handshake → spoken turn (stt → state → text → audio) → PASS. - Chunker verified by direct call:
chunk_text("สวัสดี. แล้วไง? ครับ")→["สวัสดี.", "แล้วไง?", "ครับ"]; 200-char fragment → 4×60-char hard cap. - Known bug found+fixed in review:
_SENT_ENDoriginally split on Thai vowel marks (U+0E40/41/48) — would break every word; now ASCII punctuation + ฯ (U+0E3F) only.
Status: Phase 1 complete — testable skeleton green. Next (voice-server): real STT/TTS/Opus backends, firmware protocol verification (CON-004), Cloudflare tunnel.
Phase 1.6 (2026-10-03)
- Push fixed (Cloudflare cleared):
b00312c+e91d185on origin/main. - Module-level app + 0.0.0.0 bind; live wss://zhi.moreminimore.com handshake OK.
- Device flashed (xiaozhi v2.5.0), COM3. OTA config → device next.
Phase 2 (2026-10-03) — firmware-compatible bridge + pre-flash config build
CON-003 OVERRIDE in force: build xiaozhi-esp32 v2.5.0 with ESP-IDF, set Kconfig BEFORE flash (no source rewrite).
- Verified firmware handshake from source (
main/protocols/websocket_protocol.cc,main/ota.cc):- WS auth via HTTP headers
Authorization: Bearer <token>+Device-Id(= MAC) +Protocol-Version; hello frame carries NO auth. - Server hello reply MUST have
transport:"websocket"+session_id+audio_params(ParseServerHello). - OTA: device POSTs
CONFIG_OTA_URL; response{websocket:{url,token,version}, server_time:{timestamp,timezone_offset(min)}, firmware:{version[,url]}}.server_timeis an OBJECT (not scalar). Nofirmware.url→ no forced OTA update (has_new_version_ stays false). - Firmware sends raw 20ms opus frames after hello (no per-turn JSON markers); device-side VAD.
- WS auth via HTTP headers
- Bridge: header auth (frame fallback kept), server hello
transport=websocket+session_id,/otaGET+POST returning the exact contract, config.yaml device by MAC + ota block. 21 tests pass. - ESP-IDF v6.1 installed;
export/set-target esp32s3OK. sdkconfig.defaults(appended to upstream): BOARD_TYPE_SPOTPEAR_ESP32_S3_1_54_MUMA=y, OTA_URL="https://zhi.moreminimore.com/ota", LANGUAGE_TH_TH=y.- [~] Firmware build RUNNING (proc, log esp/xiaozhi_build.log). Generated sdkconfig CONFIRMED: OTA_URL / LANGUAGE_TH_TH=y / ZH_CN unset / MUMA board all set.
- Flash built bin to COM3 (esptool) — FLASH_OK.
- Device end-to-end: OTA fetch OK → WS connect → converse.
Phase 3 (2026-10-03) — audio path wired + fresh-eyes review fixes
Pipeline in code: decode Opus → session.feed_audio → server-side VAD end-of-speech → run_turn (STT→LLM→TTS) → firmware-compatible reply (stt.text / tts sentence_start / Opus frames / tts.stop). 23 tests pass (21 + 2 regression).
Fresh-eyes review came back FAIL (2 critical + 3 major) — all fixed + regression tests added:
- C1 TTSQueue deadlock:
_tts_chunksfedq.put()with no consumer → real long reply deadlocked at 64 chunks,tts.stopnever sent. Removed queue feed from reply path (barge-in staysbarge_in()). Regression:test_long_reply_does_not_deadlock(80 sentences / 160 chunks, 5 s timeout guard). - C2 reply sample-rate mismatch: encoder hardcoded 16 kHz while TTS emits 24 kHz (board codec
AUDIO_OUTPUT_SAMPLE_RATE= 24000, firmware decodes reply Opus at codec output rate). Encoder now rate-generic (OpusEncoder, frame math from rate);_turnencodes atcfg.tts.sample_rate. No resampler needed — 24 kHz matches the codec. Regression:test_encoder_is_rate_generic(24 kHz → 480-sample/960 B frame). - M3 unbounded mic buffer (quiet room, never latches):
PCMBuffer(max_bytes=60 s)cap;Sessionbounds to 1.92 MB. - M4
tts_sentence_startper LLM token: now one per speakable sentence (tts_sentencefromchunk_text); raw tokens are display-only, not sent. - Minor 5:
_turnnow sendstts.stopon a mid-turn failure (device no longer stuck "speaking"). - Minor 6: constant-time token compare (
hmac.compare_digest). - Minor 7:
config.yamluntracked + gitignored;config.example.yamlcommitted with placeholder secrets. Note: the live token is already in remote git history — rotation is a separate owner decision (would disrupt the flashed device), NOT done here.
Next action: live voice test with the real device (mock STT/TTS, so reply text/audio is predetermined until real backends land).