Files
esp32-server/plan.md
kunthawat e96498b67e fix: resolve fresh-eyes review criticals + majors on the audio path
- C1: stop feeding the unconsumed barge-in TTSQueue in the reply path (deadlocked after 64 chunks on a real long reply)
- C2: encode reply Opus at the TTS/codec sample rate (24 kHz), not hardcoded 16 kHz; encoder now rate-generic
- M3: bound the mic PCM buffer to the 60 s max utterance
- M4: tts_sentence_start per speakable sentence, not per LLM token
- Minor: send tts.stop on a failed turn; constant-time token compare
- Sec: untrack + gitignore config.yaml; add config.example.yaml with placeholder secrets
- Tests: 23 pass (2 new regressions: long-reply deadlock, rate-generic encoder)
2026-10-03 21:55:32 +07:00

7.3 KiB
Raw Blame History

Plan — hermes-xiaozhi-bridge

Status: Phase 1.5 — live-verified (CPU host): 18 tests + live WS smoke PASS Last update: 2026-10-03 Next action: voice-server — wire real STT/TTS/Opus backends, verify protocol constants against firmware (CON-004).

Scope of this host's run

This Mac has no GPU / device / native libs, so the deliverable here is the testable bridge skeleton (AC-001..AC-005). GPU/STT/TTS/device behavior is stubbed behind interfaces and verified with mocks; the real backends are lazy imports that activate only on the voice-server.

Task breakdown

  • TASK-001 Repo, venv, deps (fastapi/websockets/pydantic/pytest…) — DEC-001
  • TASK-002 config.py + config.yaml (engines, devices, protocol constants)
  • TASK-003 xiaozhi/ — protocol (hello/hello_reply frames), messages, session, websocket endpoint
  • TASK-004 audio/ — buffer (ring), opus (lazy), VAD (energy + mock) end-of-speech
  • TASK-005 stt/ — base + typhoon + whisper (lazy) + mock
  • TASK-006 tts/ — base + jaitts (lazy) + chunker (Thai) + queue (barge-in cancel) + mock
  • TASK-007 hermes/ — client, session (multi-turn, bounded ctx), voice_profile (reasoning none)
  • TASK-008 security/auth.py — device allowlist, token hash
  • TASK-009 gpu/ — monitor, voice_models, resource_manager (voice-full/lite/off)
  • TASK-010 app/main.py — FastAPI app, /health, /ws/xiaozhi
  • TASK-011 metrics/latency.py — timestamps + End-of-Speech→First Audible Audio
  • TASK-012 tests/ — suite green (AC-001..AC-005)
  • TASK-013 README, requirements.txt, __init__.py files, .gitignore (.env.example = N/A: config.yaml is the single source of truth)

Requirement coverage

Req Task(s)
REQ-001 TASK-003, 010
REQ-002 TASK-004
REQ-003 TASK-005
REQ-004 TASK-007 (voice_profile)
REQ-005 TASK-007 (session)
REQ-006 TASK-006
REQ-007 TASK-006 (queue barge-in)
REQ-008 TASK-007 (hermes client → session_search)
REQ-009 TASK-008
REQ-010 TASK-009
REQ-011 TASK-011
REQ-012 TASK-002/008 (bind LAN; no CF logic)

Verification

  • pytest -q (CPU-only, mock backends) → expect all green.
  • python -c "import app.main" import check (AC-002).
  • Import-safety check: no GPU/opus/STT/TTS import at module load (CON-002).

Evidence

  • 2026-10-03 pytest -q → 18 passed (protocol, config, auth, chunker, barge-in queue, latency, voice-profile guard, GPU planner, VAD, e2e mock turn, app import + live WS handshake+turn).
  • 2026-10-03 python -c "import app.main" → OK; sys.modules check → no torch/typhoon/whisper/jait/opus/numpy loaded (CON-002).
  • 2026-10-03 Live boot on Windows CPU host: fixed main() ignoring config.yaml (was Config() defaults); port 8765 taken by another local service → 8766. Issued real token for xiaozhi-main (hash in config.yaml). smoke_live.py against running server: health ok, bad-token rejected, bad-device rejected, handshake → spoken turn (stt → state → text → audio) → PASS.
  • Chunker verified by direct call: chunk_text("สวัสดี. แล้วไง? ครับ") → ["สวัสดี.", "แล้วไง?", "ครับ"]; 200-char fragment → 4×60-char hard cap.
  • Known bug found+fixed in review: _SENT_END originally split on Thai vowel marks (U+0E40/41/48) — would break every word; now ASCII punctuation + ฯ (U+0E3F) only.

Status: Phase 1 complete — testable skeleton green. Next (voice-server): real STT/TTS/Opus backends, firmware protocol verification (CON-004), Cloudflare tunnel.

Phase 1.6 (2026-10-03)

  • Push fixed (Cloudflare cleared): b00312c + e91d185 on origin/main.
  • Module-level app + 0.0.0.0 bind; live wss://zhi.moreminimore.com handshake OK.
  • Device flashed (xiaozhi v2.5.0), COM3. OTA config → device next.

Phase 2 (2026-10-03) — firmware-compatible bridge + pre-flash config build

CON-003 OVERRIDE in force: build xiaozhi-esp32 v2.5.0 with ESP-IDF, set Kconfig BEFORE flash (no source rewrite).

  • Verified firmware handshake from source (main/protocols/websocket_protocol.cc, main/ota.cc):
    • WS auth via HTTP headers Authorization: Bearer <token> + Device-Id (= MAC) + Protocol-Version; hello frame carries NO auth.
    • Server hello reply MUST have transport:"websocket" + session_id + audio_params (ParseServerHello).
    • OTA: device POSTs CONFIG_OTA_URL; response {websocket:{url,token,version}, server_time:{timestamp,timezone_offset(min)}, firmware:{version[,url]}}. server_time is an OBJECT (not scalar). No firmware.url → no forced OTA update (has_new_version_ stays false).
    • Firmware sends raw 20ms opus frames after hello (no per-turn JSON markers); device-side VAD.
  • Bridge: header auth (frame fallback kept), server hello transport=websocket+session_id, /ota GET+POST returning the exact contract, config.yaml device by MAC + ota block. 21 tests pass.
  • ESP-IDF v6.1 installed; export/set-target esp32s3 OK.
  • sdkconfig.defaults (appended to upstream): BOARD_TYPE_SPOTPEAR_ESP32_S3_1_54_MUMA=y, OTA_URL="https://zhi.moreminimore.com/ota", LANGUAGE_TH_TH=y.
  • [~] Firmware build RUNNING (proc, log esp/xiaozhi_build.log). Generated sdkconfig CONFIRMED: OTA_URL / LANGUAGE_TH_TH=y / ZH_CN unset / MUMA board all set.
  • Flash built bin to COM3 (esptool) — FLASH_OK.
  • Device end-to-end: OTA fetch OK → WS connect → converse.

Phase 3 (2026-10-03) — audio path wired + fresh-eyes review fixes

Pipeline in code: decode Opus → session.feed_audio → server-side VAD end-of-speech → run_turn (STT→LLM→TTS) → firmware-compatible reply (stt.text / tts sentence_start / Opus frames / tts.stop). 23 tests pass (21 + 2 regression).

Fresh-eyes review came back FAIL (2 critical + 3 major) — all fixed + regression tests added:

  • C1 TTSQueue deadlock: _tts_chunks fed q.put() with no consumer → real long reply deadlocked at 64 chunks, tts.stop never sent. Removed queue feed from reply path (barge-in stays barge_in()). Regression: test_long_reply_does_not_deadlock (80 sentences / 160 chunks, 5 s timeout guard).
  • C2 reply sample-rate mismatch: encoder hardcoded 16 kHz while TTS emits 24 kHz (board codec AUDIO_OUTPUT_SAMPLE_RATE = 24000, firmware decodes reply Opus at codec output rate). Encoder now rate-generic (OpusEncoder, frame math from rate); _turn encodes at cfg.tts.sample_rate. No resampler needed — 24 kHz matches the codec. Regression: test_encoder_is_rate_generic (24 kHz → 480-sample/960 B frame).
  • M3 unbounded mic buffer (quiet room, never latches): PCMBuffer(max_bytes=60 s) cap; Session bounds to 1.92 MB.
  • M4 tts_sentence_start per LLM token: now one per speakable sentence (tts_sentence from chunk_text); raw tokens are display-only, not sent.
  • Minor 5: _turn now sends tts.stop on a mid-turn failure (device no longer stuck "speaking").
  • Minor 6: constant-time token compare (hmac.compare_digest).
  • Minor 7: config.yaml untracked + gitignored; config.example.yaml committed with placeholder secrets. Note: the live token is already in remote git history — rotation is a separate owner decision (would disrupt the flashed device), NOT done here.

Next action: live voice test with the real device (mock STT/TTS, so reply text/audio is predetermined until real backends land).