88 lines
5.0 KiB
Markdown
88 lines
5.0 KiB
Markdown
# Hermes-Xiaozhi Voice Companion — Project
|
||
|
||
## Purpose
|
||
Build the **Hermes-Xiaozhi Bridge**: a small glue layer that turns a Xiaozhi ESP32
|
||
into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local
|
||
Thai STT/TTS (RTX 5060 Ti). The bridge does NOT reimplement LLM/STT/TTS/agent
|
||
stacks — it adapts the Xiaozhi WebSocket protocol to internal services.
|
||
|
||
Owner intent source: `hermes-xiaozhi-voice-companion-plan.md` (attached).
|
||
|
||
## Required outcomes
|
||
- REQ-001 — Bridge accepts a Xiaozhi device over WebSocket (`/ws/xiaozhi`),
|
||
completes the documented handshake, and holds a live conversational session.
|
||
- REQ-002 — Audio in: Opus → PCM, with end-of-speech detection (VAD).
|
||
- REQ-003 — Thai STT behind an `STTEngine` interface (Typhoon primary,
|
||
faster-whisper fallback). No model hard-coded.
|
||
- REQ-004 — Hermes voice profile: reasoning **none**, tools = `session_search`
|
||
only, filesystem/shell/code-exec disabled. Read-oriented, not a worker.
|
||
- REQ-005 — Multi-turn voice session with bounded working context (~8k–16k).
|
||
- REQ-006 — Thai TTS behind a `TTSEngine` interface (JaiTTS primary),
|
||
Thai-aware chunking, streaming (first audio before full answer).
|
||
- REQ-007 — Barge-in: user speech interrupts AI speech (cancel TTS task + queue).
|
||
- REQ-008 — Cross-session recall via Hermes `session_search` (no new vector DB).
|
||
- REQ-009 — Device authentication (device_id + token_hash, allowlist) before any
|
||
session. Hermes/Qwen/vLLM never exposed to the internet directly.
|
||
- REQ-010 — GPU resource manager with manual modes: `voice-full`,
|
||
`voice-lite`, `voice-off` (auto later).
|
||
- REQ-011 — Real latency logging; headline metric = End-of-Speech → First
|
||
Audible Audio.
|
||
- REQ-012 — Cloudflare-friendly: bridge binds LAN/localhost; tunnel is external
|
||
(no Cloudflare-specific logic in the bridge).
|
||
|
||
## Constraints
|
||
- CON-001 — Python. STT/TTS ecosystem is Python-first.
|
||
- CON-002 — All backends (STT/TTS/GPU/opus) are lazy/optional imports: the
|
||
bridge must import and run (and be tested) on a machine with **no** GPU,
|
||
no device, and no native opus. Heavy deps load only when selected.
|
||
- CON-003 — No new LLM server, vector DB, memory framework, agent framework,
|
||
audio protocol, Hermes fork, firmware rewrite, or STT/TTS framework.
|
||
- CON-004 — Protocol/audio constants are **PoC assumptions** until verified
|
||
against the real Xiaozhi firmware repo. They are config-driven, not magic.
|
||
|
||
## Prohibitions
|
||
- MUST-NOT-001 — Do not hard-code an STT/TTS model name in the bridge code path.
|
||
- MUST-NOT-002 — Do not expose Hermes/Qwen/vLLM HTTP APIs to the internet.
|
||
- MUST-NOT-003 — Do not let the Voice profile write files, run shell, or execute
|
||
code (read/recall only).
|
||
|
||
## Acceptance criteria
|
||
- AC-001 — `pytest` passes on a machine with no GPU/device/native-libs
|
||
(mock backends). This is the CI-verifiable bar on this host.
|
||
- AC-002 — `python -m app.main` imports and starts a FastAPI app exposing
|
||
`/health` and `/ws/xiaozhi` (device handshake reachable).
|
||
- AC-003 — Handshake, auth, chunker, latency metrics, VAD end-of-speech,
|
||
TTS queue barge-in cancel, and GPU mode selection are each unit-tested.
|
||
- AC-004 — Voice profile asserts reasoning==none and tools=={session_search}.
|
||
- AC-005 — A full "spoken turn" runs end-to-end against mock STT/LLM/TTS in a test
|
||
(speech → text → answer text → TTS chunks → audio out) with barge-in cancel.
|
||
|
||
## Governing documents
|
||
- `plan.md` — execution plan & status (authoritative operational state).
|
||
- The attached plan (`hermes-xiaozhi-voice-companion-plan.md`) — requirements source.
|
||
|
||
## Decisions
|
||
- DEC-001 — Repo root = `/Users/kunthawat/Gitea/hermes-xiaozhi-bridge/`, Python 3.14,
|
||
venv at `.venv/`.
|
||
- DEC-002 — Protocol layer is a documented PoC (config-driven). Mark every
|
||
firmware-dependent constant with `# FIRMWARE:` so Phase 0 verification is a
|
||
grep, not an archaeology dig.
|
||
- DEC-003 — Provide `Mock*` backends (STT/LLM/TTS/GPU/opus) wired via config
|
||
`engines.*: mock` so the whole pipeline is testable on a CPU-only Mac.
|
||
|
||
## Open questions
|
||
- OQ-001 — Exact Xiaozhi firmware audio framing & hello schema (verify in firmware repo).
|
||
- OQ-002 — Hermes voice-turn API surface (local OpenAI-compatible HTTP? CLI? internal?) —
|
||
determines `app/hermes/client.py` transport.
|
||
- OQ-003 — Typhoon ASR & JaiTTS serving endpoints on the voice-server.
|
||
|
||
## Change log
|
||
- 2026-10-03 — Project initiated from attached plan. REQ/AC/CON set above.
|
||
|
||
## CON-003 OVERRIDE (2026-10-03, owner directive)
|
||
- Owner: build our own xiaozhi v2.5.0 firmware via ESP-IDF v6.1, configure before flash.
|
||
- Rationale: factory bin bakes OTA_URL=api.tenclass.net + LANGUAGE_ZH_CN; device can't reach our bridge / Thai.
|
||
- Build config: BOARD_TYPE_SPOTPEAR_ESP32_S3_1_54_MUMA, OTA_URL=https://zhi.moreminimore.com/ota, LANGUAGE_TH_TH.
|
||
- NOT a source rewrite — stock xiaozhi source, custom Kconfig defaults (OTA server + language only).
|
||
- New: bridge must serve OTA endpoint (/ota) + firmware-protocol WS auth (Authorization Bearer + Device-Id header).
|