- main() previously used Config() defaults, silently ignoring config.yaml (port + token); now loads config.yaml from repo root (BRIDGE_CONFIG override) - config.yaml: port 8766 (8765 taken on this host), real sha256 token for xiaozhi-main - smoke_live.py: live WS verification (health, auth rejection x2, handshake, full turn) - README: correct run command (python -m app.main) - HANDOFF/plan: Phase 1.5 live-verified evidence
3.6 KiB
3.6 KiB
Plan — hermes-xiaozhi-bridge
Status: Phase 1.5 — live-verified (CPU host): 18 tests + live WS smoke PASS Last update: 2026-10-03 Next action: voice-server — wire real STT/TTS/Opus backends, verify protocol constants against firmware (CON-004).
Scope of this host's run
This Mac has no GPU / device / native libs, so the deliverable here is the testable bridge skeleton (AC-001..AC-005). GPU/STT/TTS/device behavior is stubbed behind interfaces and verified with mocks; the real backends are lazy imports that activate only on the voice-server.
Task breakdown
- TASK-001 Repo, venv, deps (fastapi/websockets/pydantic/pytest…) — DEC-001
- TASK-002
config.py+config.yaml(engines, devices, protocol constants) - TASK-003
xiaozhi/— protocol (hello/hello_reply frames), messages, session, websocket endpoint - TASK-004
audio/— buffer (ring), opus (lazy), VAD (energy + mock) end-of-speech - TASK-005
stt/— base + typhoon + whisper (lazy) + mock - TASK-006
tts/— base + jaitts (lazy) + chunker (Thai) + queue (barge-in cancel) + mock - TASK-007
hermes/— client, session (multi-turn, bounded ctx), voice_profile (reasoning none) - TASK-008
security/auth.py— device allowlist, token hash - TASK-009
gpu/— monitor, voice_models, resource_manager (voice-full/lite/off) - TASK-010
app/main.py— FastAPI app,/health,/ws/xiaozhi - TASK-011
metrics/latency.py— timestamps + End-of-Speech→First Audible Audio - TASK-012
tests/— suite green (AC-001..AC-005) - TASK-013 README, requirements.txt,
__init__.pyfiles, .gitignore (.env.example= N/A: config.yaml is the single source of truth)
Requirement coverage
| Req | Task(s) |
|---|---|
| REQ-001 | TASK-003, 010 |
| REQ-002 | TASK-004 |
| REQ-003 | TASK-005 |
| REQ-004 | TASK-007 (voice_profile) |
| REQ-005 | TASK-007 (session) |
| REQ-006 | TASK-006 |
| REQ-007 | TASK-006 (queue barge-in) |
| REQ-008 | TASK-007 (hermes client → session_search) |
| REQ-009 | TASK-008 |
| REQ-010 | TASK-009 |
| REQ-011 | TASK-011 |
| REQ-012 | TASK-002/008 (bind LAN; no CF logic) |
Verification
pytest -q(CPU-only, mock backends) → expect all green.python -c "import app.main"import check (AC-002).- Import-safety check: no GPU/opus/STT/TTS import at module load (CON-002).
Evidence
- 2026-10-03
pytest -q→ 18 passed (protocol, config, auth, chunker, barge-in queue, latency, voice-profile guard, GPU planner, VAD, e2e mock turn, app import + live WS handshake+turn). - 2026-10-03
python -c "import app.main"→ OK;sys.modulescheck → no torch/typhoon/whisper/jait/opus/numpy loaded (CON-002). - 2026-10-03 Live boot on Windows CPU host: fixed
main()ignoring config.yaml (wasConfig()defaults); port 8765 taken by another local service → 8766. Issued real token forxiaozhi-main(hash in config.yaml).smoke_live.pyagainst running server: health ok, bad-token rejected, bad-device rejected, handshake → spoken turn (stt → state → text → audio) → PASS. - Chunker verified by direct call:
chunk_text("สวัสดี. แล้วไง? ครับ")→["สวัสดี.", "แล้วไง?", "ครับ"]; 200-char fragment → 4×60-char hard cap. - Known bug found+fixed in review:
_SENT_ENDoriginally split on Thai vowel marks (U+0E40/41/48) — would break every word; now ASCII punctuation + ฯ (U+0E3F) only.
Status: Phase 1 complete — testable skeleton green. Next (voice-server): real STT/TTS/Opus backends, firmware protocol verification (CON-004), Cloudflare tunnel.