5.0 KiB
5.0 KiB
Hermes-Xiaozhi Voice Companion — Project
Purpose
Build the Hermes-Xiaozhi Bridge: a small glue layer that turns a Xiaozhi ESP32 into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local Thai STT/TTS (RTX 5060 Ti). The bridge does NOT reimplement LLM/STT/TTS/agent stacks — it adapts the Xiaozhi WebSocket protocol to internal services.
Owner intent source: hermes-xiaozhi-voice-companion-plan.md (attached).
Required outcomes
- REQ-001 — Bridge accepts a Xiaozhi device over WebSocket (
/ws/xiaozhi), completes the documented handshake, and holds a live conversational session. - REQ-002 — Audio in: Opus → PCM, with end-of-speech detection (VAD).
- REQ-003 — Thai STT behind an
STTEngineinterface (Typhoon primary, faster-whisper fallback). No model hard-coded. - REQ-004 — Hermes voice profile: reasoning none, tools =
session_searchonly, filesystem/shell/code-exec disabled. Read-oriented, not a worker. - REQ-005 — Multi-turn voice session with bounded working context (~8k–16k).
- REQ-006 — Thai TTS behind a
TTSEngineinterface (JaiTTS primary), Thai-aware chunking, streaming (first audio before full answer). - REQ-007 — Barge-in: user speech interrupts AI speech (cancel TTS task + queue).
- REQ-008 — Cross-session recall via Hermes
session_search(no new vector DB). - REQ-009 — Device authentication (device_id + token_hash, allowlist) before any session. Hermes/Qwen/vLLM never exposed to the internet directly.
- REQ-010 — GPU resource manager with manual modes:
voice-full,voice-lite,voice-off(auto later). - REQ-011 — Real latency logging; headline metric = End-of-Speech → First Audible Audio.
- REQ-012 — Cloudflare-friendly: bridge binds LAN/localhost; tunnel is external (no Cloudflare-specific logic in the bridge).
Constraints
- CON-001 — Python. STT/TTS ecosystem is Python-first.
- CON-002 — All backends (STT/TTS/GPU/opus) are lazy/optional imports: the bridge must import and run (and be tested) on a machine with no GPU, no device, and no native opus. Heavy deps load only when selected.
- CON-003 — No new LLM server, vector DB, memory framework, agent framework, audio protocol, Hermes fork, firmware rewrite, or STT/TTS framework.
- CON-004 — Protocol/audio constants are PoC assumptions until verified against the real Xiaozhi firmware repo. They are config-driven, not magic.
Prohibitions
- MUST-NOT-001 — Do not hard-code an STT/TTS model name in the bridge code path.
- MUST-NOT-002 — Do not expose Hermes/Qwen/vLLM HTTP APIs to the internet.
- MUST-NOT-003 — Do not let the Voice profile write files, run shell, or execute code (read/recall only).
Acceptance criteria
- AC-001 —
pytestpasses on a machine with no GPU/device/native-libs (mock backends). This is the CI-verifiable bar on this host. - AC-002 —
python -m app.mainimports and starts a FastAPI app exposing/healthand/ws/xiaozhi(device handshake reachable). - AC-003 — Handshake, auth, chunker, latency metrics, VAD end-of-speech, TTS queue barge-in cancel, and GPU mode selection are each unit-tested.
- AC-004 — Voice profile asserts reasoning==none and tools=={session_search}.
- AC-005 — A full "spoken turn" runs end-to-end against mock STT/LLM/TTS in a test (speech → text → answer text → TTS chunks → audio out) with barge-in cancel.
Governing documents
plan.md— execution plan & status (authoritative operational state).- The attached plan (
hermes-xiaozhi-voice-companion-plan.md) — requirements source.
Decisions
- DEC-001 — Repo root =
/Users/kunthawat/Gitea/hermes-xiaozhi-bridge/, Python 3.14, venv at.venv/. - DEC-002 — Protocol layer is a documented PoC (config-driven). Mark every
firmware-dependent constant with
# FIRMWARE:so Phase 0 verification is a grep, not an archaeology dig. - DEC-003 — Provide
Mock*backends (STT/LLM/TTS/GPU/opus) wired via configengines.*: mockso the whole pipeline is testable on a CPU-only Mac.
Open questions
- OQ-001 — Exact Xiaozhi firmware audio framing & hello schema (verify in firmware repo).
- OQ-002 — Hermes voice-turn API surface (local OpenAI-compatible HTTP? CLI? internal?) —
determines
app/hermes/client.pytransport. - OQ-003 — Typhoon ASR & JaiTTS serving endpoints on the voice-server.
Change log
- 2026-10-03 — Project initiated from attached plan. REQ/AC/CON set above.
CON-003 OVERRIDE (2026-10-03, owner directive)
- Owner: build our own xiaozhi v2.5.0 firmware via ESP-IDF v6.1, configure before flash.
- Rationale: factory bin bakes OTA_URL=api.tenclass.net + LANGUAGE_ZH_CN; device can't reach our bridge / Thai.
- Build config: BOARD_TYPE_SPOTPEAR_ESP32_S3_1_54_MUMA, OTA_URL=https://zhi.moreminimore.com/ota, LANGUAGE_TH_TH.
- NOT a source rewrite — stock xiaozhi source, custom Kconfig defaults (OTA server + language only).
- New: bridge must serve OTA endpoint (/ota) + firmware-protocol WS auth (Authorization Bearer + Device-Id header).