662c663cda558dff0b67794b8b233fafd4f31e20
hermes-xiaozhi-bridge
Glue layer that turns a Xiaozhi ESP32 into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local Thai STT/TTS (RTX 5060 Ti).
The bridge does not reimplement the LLM, STT, TTS, or audio protocol — it adapts the Xiaozhi WebSocket protocol to services inside the machine.
Owner intent & plan:
hermes-xiaozhi-voice-companion-plan.md(attached copy). Operational status:plan.md. Requirements:project.md.
What exists (Phase 1 — testable skeleton)
app/main.py— FastAPI app,/ws/xiaozhiendpoint,/healthapp/xiaozhi/— WebSocket loop, protocol (hello/hello-reply), messages, session state machineapp/audio/— PCM buffer, VAD (energy + Silero stub), Opus (passthrough + real, lazy)app/stt/— interface + mock + Typhoon + faster-whisper (lazy)app/tts/— interface + mock + JaiTTS (lazy), Thai chunker, barge-in queueapp/hermes/— voice client (mock transport now, OpenAI-HTTP later), voice profile guardapp/security/— device allowlist + token hashapp/metrics/— latency logger (EoS → first-audio stages)app/gpu/— VRAM monitor + resource planner (FULL/LITE/OFF)
All backends default to mock in config.yaml, so the bridge runs and tests
pass on any machine — no GPU, no device, no native libs.
Run
cd hermes-xiaozhi-bridge
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python -m app.main
# (loads config.yaml from the repo root; override with BRIDGE_CONFIG=/path)
# equivalent: .venv/bin/uvicorn app.main:app --host 0.0.0.0 --port 8766
.venv/bin/python -m pytest -q
Voice profile (REQ-004)
The bridge asserts at startup: reasoning = none, tools = session_search
only, no filesystem/shell/code-exec. Bad config fails boot, not conversation.
Not yet
- Real STT/TTS/Opus backends (lazy, need voice-server GPU)
- Firmware-verified protocol constants (marked
# FIRMWARE) - Cloudflare tunnel config (external to the bridge)
- MCP device bridge (Phase 10, after voice MVP)
Description
Languages
Python
100%