# Hermes-Xiaozhi Voice Companion — Project ## Purpose Build the **Hermes-Xiaozhi Bridge**: a small glue layer that turns a Xiaozhi ESP32 into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local Thai STT/TTS (RTX 5060 Ti). The bridge does NOT reimplement LLM/STT/TTS/agent stacks — it adapts the Xiaozhi WebSocket protocol to internal services. Owner intent source: `hermes-xiaozhi-voice-companion-plan.md` (attached). ## Required outcomes - REQ-001 — Bridge accepts a Xiaozhi device over WebSocket (`/ws/xiaozhi`), completes the documented handshake, and holds a live conversational session. - REQ-002 — Audio in: Opus → PCM, with end-of-speech detection (VAD). - REQ-003 — Thai STT behind an `STTEngine` interface (Typhoon primary, faster-whisper fallback). No model hard-coded. - REQ-004 — Hermes voice profile: reasoning **none**, tools = `session_search` only, filesystem/shell/code-exec disabled. Read-oriented, not a worker. - REQ-005 — Multi-turn voice session with bounded working context (~8k–16k). - REQ-006 — Thai TTS behind a `TTSEngine` interface (JaiTTS primary), Thai-aware chunking, streaming (first audio before full answer). - REQ-007 — Barge-in: user speech interrupts AI speech (cancel TTS task + queue). - REQ-008 — Cross-session recall via Hermes `session_search` (no new vector DB). - REQ-009 — Device authentication (device_id + token_hash, allowlist) before any session. Hermes/Qwen/vLLM never exposed to the internet directly. - REQ-010 — GPU resource manager with manual modes: `voice-full`, `voice-lite`, `voice-off` (auto later). - REQ-011 — Real latency logging; headline metric = End-of-Speech → First Audible Audio. - REQ-012 — Cloudflare-friendly: bridge binds LAN/localhost; tunnel is external (no Cloudflare-specific logic in the bridge). ## Constraints - CON-001 — Python. STT/TTS ecosystem is Python-first. - CON-002 — All backends (STT/TTS/GPU/opus) are lazy/optional imports: the bridge must import and run (and be tested) on a machine with **no** GPU, no device, and no native opus. Heavy deps load only when selected. - CON-003 — No new LLM server, vector DB, memory framework, agent framework, audio protocol, Hermes fork, firmware rewrite, or STT/TTS framework. - CON-004 — Protocol/audio constants are **PoC assumptions** until verified against the real Xiaozhi firmware repo. They are config-driven, not magic. ## Prohibitions - MUST-NOT-001 — Do not hard-code an STT/TTS model name in the bridge code path. - MUST-NOT-002 — Do not expose Hermes/Qwen/vLLM HTTP APIs to the internet. - MUST-NOT-003 — Do not let the Voice profile write files, run shell, or execute code (read/recall only). ## Acceptance criteria - AC-001 — `pytest` passes on a machine with no GPU/device/native-libs (mock backends). This is the CI-verifiable bar on this host. - AC-002 — `python -m app.main` imports and starts a FastAPI app exposing `/health` and `/ws/xiaozhi` (device handshake reachable). - AC-003 — Handshake, auth, chunker, latency metrics, VAD end-of-speech, TTS queue barge-in cancel, and GPU mode selection are each unit-tested. - AC-004 — Voice profile asserts reasoning==none and tools=={session_search}. - AC-005 — A full "spoken turn" runs end-to-end against mock STT/LLM/TTS in a test (speech → text → answer text → TTS chunks → audio out) with barge-in cancel. ## Governing documents - `plan.md` — execution plan & status (authoritative operational state). - The attached plan (`hermes-xiaozhi-voice-companion-plan.md`) — requirements source. ## Decisions - DEC-001 — Repo root = `/Users/kunthawat/Gitea/hermes-xiaozhi-bridge/`, Python 3.14, venv at `.venv/`. - DEC-002 — Protocol layer is a documented PoC (config-driven). Mark every firmware-dependent constant with `# FIRMWARE:` so Phase 0 verification is a grep, not an archaeology dig. - DEC-003 — Provide `Mock*` backends (STT/LLM/TTS/GPU/opus) wired via config `engines.*: mock` so the whole pipeline is testable on a CPU-only Mac. ## Open questions - OQ-001 — Exact Xiaozhi firmware audio framing & hello schema (verify in firmware repo). - OQ-002 — Hermes voice-turn API surface (local OpenAI-compatible HTTP? CLI? internal?) — determines `app/hermes/client.py` transport. - OQ-003 — Typhoon ASR & JaiTTS serving endpoints on the voice-server. ## Change log - 2026-10-03 — Project initiated from attached plan. REQ/AC/CON set above. ## CON-003 OVERRIDE (2026-10-03, owner directive) - Owner: build our own xiaozhi v2.5.0 firmware via ESP-IDF v6.1, configure before flash. - Rationale: factory bin bakes OTA_URL=api.tenclass.net + LANGUAGE_ZH_CN; device can't reach our bridge / Thai. - Build config: BOARD_TYPE_SPOTPEAR_ESP32_S3_1_54_MUMA, OTA_URL=https://zhi.moreminimore.com/ota, LANGUAGE_TH_TH. - NOT a source rewrite — stock xiaozhi source, custom Kconfig defaults (OTA server + language only). - New: bridge must serve OTA endpoint (/ota) + firmware-protocol WS auth (Authorization Bearer + Device-Id header).