Files
esp32-server/project.md

5.0 KiB
Raw Permalink Blame History

Hermes-Xiaozhi Voice Companion — Project

Purpose

Build the Hermes-Xiaozhi Bridge: a small glue layer that turns a Xiaozhi ESP32 into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local Thai STT/TTS (RTX 5060 Ti). The bridge does NOT reimplement LLM/STT/TTS/agent stacks — it adapts the Xiaozhi WebSocket protocol to internal services.

Owner intent source: hermes-xiaozhi-voice-companion-plan.md (attached).

Required outcomes

  • REQ-001 — Bridge accepts a Xiaozhi device over WebSocket (/ws/xiaozhi), completes the documented handshake, and holds a live conversational session.
  • REQ-002 — Audio in: Opus → PCM, with end-of-speech detection (VAD).
  • REQ-003 — Thai STT behind an STTEngine interface (Typhoon primary, faster-whisper fallback). No model hard-coded.
  • REQ-004 — Hermes voice profile: reasoning none, tools = session_search only, filesystem/shell/code-exec disabled. Read-oriented, not a worker.
  • REQ-005 — Multi-turn voice session with bounded working context (~8k–16k).
  • REQ-006 — Thai TTS behind a TTSEngine interface (JaiTTS primary), Thai-aware chunking, streaming (first audio before full answer).
  • REQ-007 — Barge-in: user speech interrupts AI speech (cancel TTS task + queue).
  • REQ-008 — Cross-session recall via Hermes session_search (no new vector DB).
  • REQ-009 — Device authentication (device_id + token_hash, allowlist) before any session. Hermes/Qwen/vLLM never exposed to the internet directly.
  • REQ-010 — GPU resource manager with manual modes: voice-full, voice-lite, voice-off (auto later).
  • REQ-011 — Real latency logging; headline metric = End-of-Speech → First Audible Audio.
  • REQ-012 — Cloudflare-friendly: bridge binds LAN/localhost; tunnel is external (no Cloudflare-specific logic in the bridge).

Constraints

  • CON-001 — Python. STT/TTS ecosystem is Python-first.
  • CON-002 — All backends (STT/TTS/GPU/opus) are lazy/optional imports: the bridge must import and run (and be tested) on a machine with no GPU, no device, and no native opus. Heavy deps load only when selected.
  • CON-003 — No new LLM server, vector DB, memory framework, agent framework, audio protocol, Hermes fork, firmware rewrite, or STT/TTS framework.
  • CON-004 — Protocol/audio constants are PoC assumptions until verified against the real Xiaozhi firmware repo. They are config-driven, not magic.

Prohibitions

  • MUST-NOT-001 — Do not hard-code an STT/TTS model name in the bridge code path.
  • MUST-NOT-002 — Do not expose Hermes/Qwen/vLLM HTTP APIs to the internet.
  • MUST-NOT-003 — Do not let the Voice profile write files, run shell, or execute code (read/recall only).

Acceptance criteria

  • AC-001 — pytest passes on a machine with no GPU/device/native-libs (mock backends). This is the CI-verifiable bar on this host.
  • AC-002 — python -m app.main imports and starts a FastAPI app exposing /health and /ws/xiaozhi (device handshake reachable).
  • AC-003 — Handshake, auth, chunker, latency metrics, VAD end-of-speech, TTS queue barge-in cancel, and GPU mode selection are each unit-tested.
  • AC-004 — Voice profile asserts reasoning==none and tools=={session_search}.
  • AC-005 — A full "spoken turn" runs end-to-end against mock STT/LLM/TTS in a test (speech → text → answer text → TTS chunks → audio out) with barge-in cancel.

Governing documents

  • plan.md — execution plan & status (authoritative operational state).
  • The attached plan (hermes-xiaozhi-voice-companion-plan.md) — requirements source.

Decisions

  • DEC-001 — Repo root = /Users/kunthawat/Gitea/hermes-xiaozhi-bridge/, Python 3.14, venv at .venv/.
  • DEC-002 — Protocol layer is a documented PoC (config-driven). Mark every firmware-dependent constant with # FIRMWARE: so Phase 0 verification is a grep, not an archaeology dig.
  • DEC-003 — Provide Mock* backends (STT/LLM/TTS/GPU/opus) wired via config engines.*: mock so the whole pipeline is testable on a CPU-only Mac.

Open questions

  • OQ-001 — Exact Xiaozhi firmware audio framing & hello schema (verify in firmware repo).
  • OQ-002 — Hermes voice-turn API surface (local OpenAI-compatible HTTP? CLI? internal?) — determines app/hermes/client.py transport.
  • OQ-003 — Typhoon ASR & JaiTTS serving endpoints on the voice-server.

Change log

  • 2026-10-03 — Project initiated from attached plan. REQ/AC/CON set above.

CON-003 OVERRIDE (2026-10-03, owner directive)

  • Owner: build our own xiaozhi v2.5.0 firmware via ESP-IDF v6.1, configure before flash.
  • Rationale: factory bin bakes OTA_URL=api.tenclass.net + LANGUAGE_ZH_CN; device can't reach our bridge / Thai.
  • Build config: BOARD_TYPE_SPOTPEAR_ESP32_S3_1_54_MUMA, OTA_URL=https://zhi.moreminimore.com/ota, LANGUAGE_TH_TH.
  • NOT a source rewrite — stock xiaozhi source, custom Kconfig defaults (OTA server + language only).
  • New: bridge must serve OTA endpoint (/ota) + firmware-protocol WS auth (Authorization Bearer + Device-Id header).