Files
esp32-server/plan.md

3.8 KiB
Raw Blame History

Plan — hermes-xiaozhi-bridge

Status: Phase 1.5 — live-verified (CPU host): 18 tests + live WS smoke PASS Last update: 2026-10-03 Next action: voice-server — wire real STT/TTS/Opus backends, verify protocol constants against firmware (CON-004).

Scope of this host's run

This Mac has no GPU / device / native libs, so the deliverable here is the testable bridge skeleton (AC-001..AC-005). GPU/STT/TTS/device behavior is stubbed behind interfaces and verified with mocks; the real backends are lazy imports that activate only on the voice-server.

Task breakdown

  • TASK-001 Repo, venv, deps (fastapi/websockets/pydantic/pytest…) — DEC-001
  • TASK-002 config.py + config.yaml (engines, devices, protocol constants)
  • TASK-003 xiaozhi/ — protocol (hello/hello_reply frames), messages, session, websocket endpoint
  • TASK-004 audio/ — buffer (ring), opus (lazy), VAD (energy + mock) end-of-speech
  • TASK-005 stt/ — base + typhoon + whisper (lazy) + mock
  • TASK-006 tts/ — base + jaitts (lazy) + chunker (Thai) + queue (barge-in cancel) + mock
  • TASK-007 hermes/ — client, session (multi-turn, bounded ctx), voice_profile (reasoning none)
  • TASK-008 security/auth.py — device allowlist, token hash
  • TASK-009 gpu/ — monitor, voice_models, resource_manager (voice-full/lite/off)
  • TASK-010 app/main.py — FastAPI app, /health, /ws/xiaozhi
  • TASK-011 metrics/latency.py — timestamps + End-of-Speech→First Audible Audio
  • TASK-012 tests/ — suite green (AC-001..AC-005)
  • TASK-013 README, requirements.txt, __init__.py files, .gitignore (.env.example = N/A: config.yaml is the single source of truth)

Requirement coverage

Req Task(s)
REQ-001 TASK-003, 010
REQ-002 TASK-004
REQ-003 TASK-005
REQ-004 TASK-007 (voice_profile)
REQ-005 TASK-007 (session)
REQ-006 TASK-006
REQ-007 TASK-006 (queue barge-in)
REQ-008 TASK-007 (hermes client → session_search)
REQ-009 TASK-008
REQ-010 TASK-009
REQ-011 TASK-011
REQ-012 TASK-002/008 (bind LAN; no CF logic)

Verification

  • pytest -q (CPU-only, mock backends) → expect all green.
  • python -c "import app.main" import check (AC-002).
  • Import-safety check: no GPU/opus/STT/TTS import at module load (CON-002).

Evidence

  • 2026-10-03 pytest -q → 18 passed (protocol, config, auth, chunker, barge-in queue, latency, voice-profile guard, GPU planner, VAD, e2e mock turn, app import + live WS handshake+turn).
  • 2026-10-03 python -c "import app.main" → OK; sys.modules check → no torch/typhoon/whisper/jait/opus/numpy loaded (CON-002).
  • 2026-10-03 Live boot on Windows CPU host: fixed main() ignoring config.yaml (was Config() defaults); port 8765 taken by another local service → 8766. Issued real token for xiaozhi-main (hash in config.yaml). smoke_live.py against running server: health ok, bad-token rejected, bad-device rejected, handshake → spoken turn (stt → state → text → audio) → PASS.
  • Chunker verified by direct call: chunk_text("สวัสดี. แล้วไง? ครับ") → ["สวัสดี.", "แล้วไง?", "ครับ"]; 200-char fragment → 4×60-char hard cap.
  • Known bug found+fixed in review: _SENT_END originally split on Thai vowel marks (U+0E40/41/48) — would break every word; now ASCII punctuation + ฯ (U+0E3F) only.

Status: Phase 1 complete — testable skeleton green. Next (voice-server): real STT/TTS/Opus backends, firmware protocol verification (CON-004), Cloudflare tunnel.

Phase 1.6 (2026-10-03)

  • Push fixed (Cloudflare cleared): b00312c + e91d185 on origin/main.
  • Module-level app + 0.0.0.0 bind; live wss://zhi.moreminimore.com handshake OK.
  • Device flashed (xiaozhi v2.5.0), COM3. OTA config → device next.