Files
esp32-server/README.md

2.1 KiB

hermes-xiaozhi-bridge

Glue layer that turns a Xiaozhi ESP32 into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local Thai STT/TTS (RTX 5060 Ti).

The bridge does not reimplement the LLM, STT, TTS, or audio protocol — it adapts the Xiaozhi WebSocket protocol to services inside the machine.

Owner intent & plan: hermes-xiaozhi-voice-companion-plan.md (attached copy). Operational status: plan.md. Requirements: project.md.

What exists (Phase 1 — testable skeleton)

  • app/main.py — FastAPI app, /ws/xiaozhi endpoint, /health
  • app/xiaozhi/ — WebSocket loop, protocol (hello/hello-reply), messages, session state machine
  • app/audio/ — PCM buffer, VAD (energy + Silero stub), Opus (passthrough + real, lazy)
  • app/stt/ — interface + mock + Typhoon + faster-whisper (lazy)
  • app/tts/ — interface + mock + JaiTTS (lazy), Thai chunker, barge-in queue
  • app/hermes/ — voice client (mock transport now, OpenAI-HTTP later), voice profile guard
  • app/security/ — device allowlist + token hash
  • app/metrics/ — latency logger (EoS → first-audio stages)
  • app/gpu/ — VRAM monitor + resource planner (FULL/LITE/OFF)

All backends default to mock in config.yaml, so the bridge runs and tests pass on any machine — no GPU, no device, no native libs.

Run

cd hermes-xiaozhi-bridge
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python -m app.main
# (loads config.yaml from the repo root; override with BRIDGE_CONFIG=/path)
# equivalent: .venv/bin/uvicorn app.main:app --host 0.0.0.0 --port 8766
.venv/bin/python -m pytest -q

Voice profile (REQ-004)

The bridge asserts at startup: reasoning = none, tools = session_search only, no filesystem/shell/code-exec. Bad config fails boot, not conversation.

Not yet

  • Real STT/TTS/Opus backends (lazy, need voice-server GPU)
  • Firmware-verified protocol constants (marked # FIRMWARE)
  • Cloudflare tunnel config (external to the bridge)
  • MCP device bridge (Phase 10, after voice MVP)