Files
esp32-server/project.md

88 lines
5.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Hermes-Xiaozhi Voice Companion — Project
## Purpose
Build the **Hermes-Xiaozhi Bridge**: a small glue layer that turns a Xiaozhi ESP32
into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local
Thai STT/TTS (RTX 5060 Ti). The bridge does NOT reimplement LLM/STT/TTS/agent
stacks — it adapts the Xiaozhi WebSocket protocol to internal services.
Owner intent source: `hermes-xiaozhi-voice-companion-plan.md` (attached).
## Required outcomes
- REQ-001 — Bridge accepts a Xiaozhi device over WebSocket (`/ws/xiaozhi`),
completes the documented handshake, and holds a live conversational session.
- REQ-002 — Audio in: Opus → PCM, with end-of-speech detection (VAD).
- REQ-003 — Thai STT behind an `STTEngine` interface (Typhoon primary,
faster-whisper fallback). No model hard-coded.
- REQ-004 — Hermes voice profile: reasoning **none**, tools = `session_search`
only, filesystem/shell/code-exec disabled. Read-oriented, not a worker.
- REQ-005 — Multi-turn voice session with bounded working context (~8k–16k).
- REQ-006 — Thai TTS behind a `TTSEngine` interface (JaiTTS primary),
Thai-aware chunking, streaming (first audio before full answer).
- REQ-007 — Barge-in: user speech interrupts AI speech (cancel TTS task + queue).
- REQ-008 — Cross-session recall via Hermes `session_search` (no new vector DB).
- REQ-009 — Device authentication (device_id + token_hash, allowlist) before any
session. Hermes/Qwen/vLLM never exposed to the internet directly.
- REQ-010 — GPU resource manager with manual modes: `voice-full`,
`voice-lite`, `voice-off` (auto later).
- REQ-011 — Real latency logging; headline metric = End-of-Speech → First
Audible Audio.
- REQ-012 — Cloudflare-friendly: bridge binds LAN/localhost; tunnel is external
(no Cloudflare-specific logic in the bridge).
## Constraints
- CON-001 — Python. STT/TTS ecosystem is Python-first.
- CON-002 — All backends (STT/TTS/GPU/opus) are lazy/optional imports: the
bridge must import and run (and be tested) on a machine with **no** GPU,
no device, and no native opus. Heavy deps load only when selected.
- CON-003 — No new LLM server, vector DB, memory framework, agent framework,
audio protocol, Hermes fork, firmware rewrite, or STT/TTS framework.
- CON-004 — Protocol/audio constants are **PoC assumptions** until verified
against the real Xiaozhi firmware repo. They are config-driven, not magic.
## Prohibitions
- MUST-NOT-001 — Do not hard-code an STT/TTS model name in the bridge code path.
- MUST-NOT-002 — Do not expose Hermes/Qwen/vLLM HTTP APIs to the internet.
- MUST-NOT-003 — Do not let the Voice profile write files, run shell, or execute
code (read/recall only).
## Acceptance criteria
- AC-001 — `pytest` passes on a machine with no GPU/device/native-libs
(mock backends). This is the CI-verifiable bar on this host.
- AC-002 — `python -m app.main` imports and starts a FastAPI app exposing
`/health` and `/ws/xiaozhi` (device handshake reachable).
- AC-003 — Handshake, auth, chunker, latency metrics, VAD end-of-speech,
TTS queue barge-in cancel, and GPU mode selection are each unit-tested.
- AC-004 — Voice profile asserts reasoning==none and tools=={session_search}.
- AC-005 — A full "spoken turn" runs end-to-end against mock STT/LLM/TTS in a test
(speech → text → answer text → TTS chunks → audio out) with barge-in cancel.
## Governing documents
- `plan.md` — execution plan & status (authoritative operational state).
- The attached plan (`hermes-xiaozhi-voice-companion-plan.md`) — requirements source.
## Decisions
- DEC-001 — Repo root = `/Users/kunthawat/Gitea/hermes-xiaozhi-bridge/`, Python 3.14,
venv at `.venv/`.
- DEC-002 — Protocol layer is a documented PoC (config-driven). Mark every
firmware-dependent constant with `# FIRMWARE:` so Phase 0 verification is a
grep, not an archaeology dig.
- DEC-003 — Provide `Mock*` backends (STT/LLM/TTS/GPU/opus) wired via config
`engines.*: mock` so the whole pipeline is testable on a CPU-only Mac.
## Open questions
- OQ-001 — Exact Xiaozhi firmware audio framing & hello schema (verify in firmware repo).
- OQ-002 — Hermes voice-turn API surface (local OpenAI-compatible HTTP? CLI? internal?) —
determines `app/hermes/client.py` transport.
- OQ-003 — Typhoon ASR & JaiTTS serving endpoints on the voice-server.
## Change log
- 2026-10-03 — Project initiated from attached plan. REQ/AC/CON set above.
## CON-003 OVERRIDE (2026-10-03, owner directive)
- Owner: build our own xiaozhi v2.5.0 firmware via ESP-IDF v6.1, configure before flash.
- Rationale: factory bin bakes OTA_URL=api.tenclass.net + LANGUAGE_ZH_CN; device can't reach our bridge / Thai.
- Build config: BOARD_TYPE_SPOTPEAR_ESP32_S3_1_54_MUMA, OTA_URL=https://zhi.moreminimore.com/ota, LANGUAGE_TH_TH.
- NOT a source rewrite — stock xiaozhi source, custom Kconfig defaults (OTA server + language only).
- New: bridge must serve OTA endpoint (/ota) + firmware-protocol WS auth (Authorization Bearer + Device-Id header).