- main() previously used Config() defaults, silently ignoring config.yaml (port + token); now loads config.yaml from repo root (BRIDGE_CONFIG override) - config.yaml: port 8766 (8765 taken on this host), real sha256 token for xiaozhi-main - smoke_live.py: live WS verification (health, auth rejection x2, handshake, full turn) - README: correct run command (python -m app.main) - HANDOFF/plan: Phase 1.5 live-verified evidence
2.0 KiB
2.0 KiB
hermes-xiaozhi-bridge
Glue layer that turns a Xiaozhi ESP32 into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local Thai STT/TTS (RTX 5060 Ti).
The bridge does not reimplement the LLM, STT, TTS, or audio protocol — it adapts the Xiaozhi WebSocket protocol to services inside the machine.
Owner intent & plan:
hermes-xiaozhi-voice-companion-plan.md(attached copy). Operational status:plan.md. Requirements:project.md.
What exists (Phase 1 — testable skeleton)
app/main.py— FastAPI app,/ws/xiaozhiendpoint,/healthapp/xiaozhi/— WebSocket loop, protocol (hello/hello-reply), messages, session state machineapp/audio/— PCM buffer, VAD (energy + Silero stub), Opus (passthrough + real, lazy)app/stt/— interface + mock + Typhoon + faster-whisper (lazy)app/tts/— interface + mock + JaiTTS (lazy), Thai chunker, barge-in queueapp/hermes/— voice client (mock transport now, OpenAI-HTTP later), voice profile guardapp/security/— device allowlist + token hashapp/metrics/— latency logger (EoS → first-audio stages)app/gpu/— VRAM monitor + resource planner (FULL/LITE/OFF)
All backends default to mock in config.yaml, so the bridge runs and tests
pass on any machine — no GPU, no device, no native libs.
Run
cd hermes-xiaozhi-bridge
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python -m app.main
# (loads config.yaml from the repo root; override with BRIDGE_CONFIG=/path)
.venv/bin/python -m pytest -q
Voice profile (REQ-004)
The bridge asserts at startup: reasoning = none, tools = session_search
only, no filesystem/shell/code-exec. Bad config fails boot, not conversation.
Not yet
- Real STT/TTS/Opus backends (lazy, need voice-server GPU)
- Firmware-verified protocol constants (marked
# FIRMWARE) - Cloudflare tunnel config (external to the bridge)
- MCP device bridge (Phase 10, after voice MVP)