kunthawat e96498b67e fix: resolve fresh-eyes review criticals + majors on the audio path
- C1: stop feeding the unconsumed barge-in TTSQueue in the reply path (deadlocked after 64 chunks on a real long reply)
- C2: encode reply Opus at the TTS/codec sample rate (24 kHz), not hardcoded 16 kHz; encoder now rate-generic
- M3: bound the mic PCM buffer to the 60 s max utterance
- M4: tts_sentence_start per speakable sentence, not per LLM token
- Minor: send tts.stop on a failed turn; constant-time token compare
- Sec: untrack + gitignore config.yaml; add config.example.yaml with placeholder secrets
- Tests: 23 pass (2 new regressions: long-reply deadlock, rate-generic encoder)
2026-10-03 21:55:32 +07:00

hermes-xiaozhi-bridge

Glue layer that turns a Xiaozhi ESP32 into a Thai voice companion backed by local Hermes + Qwen 3.8 (V100) and local Thai STT/TTS (RTX 5060 Ti).

The bridge does not reimplement the LLM, STT, TTS, or audio protocol — it adapts the Xiaozhi WebSocket protocol to services inside the machine.

Owner intent & plan: hermes-xiaozhi-voice-companion-plan.md (attached copy). Operational status: plan.md. Requirements: project.md.

What exists (Phase 1 — testable skeleton)

  • app/main.py — FastAPI app, /ws/xiaozhi endpoint, /health
  • app/xiaozhi/ — WebSocket loop, protocol (hello/hello-reply), messages, session state machine
  • app/audio/ — PCM buffer, VAD (energy + Silero stub), Opus (passthrough + real, lazy)
  • app/stt/ — interface + mock + Typhoon + faster-whisper (lazy)
  • app/tts/ — interface + mock + JaiTTS (lazy), Thai chunker, barge-in queue
  • app/hermes/ — voice client (mock transport now, OpenAI-HTTP later), voice profile guard
  • app/security/ — device allowlist + token hash
  • app/metrics/ — latency logger (EoS → first-audio stages)
  • app/gpu/ — VRAM monitor + resource planner (FULL/LITE/OFF)

All backends default to mock in config.yaml, so the bridge runs and tests pass on any machine — no GPU, no device, no native libs.

Run

cd hermes-xiaozhi-bridge
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python -m app.main
# (loads config.yaml from the repo root; override with BRIDGE_CONFIG=/path)
# equivalent: .venv/bin/uvicorn app.main:app --host 0.0.0.0 --port 8766
.venv/bin/python -m pytest -q

Voice profile (REQ-004)

The bridge asserts at startup: reasoning = none, tools = session_search only, no filesystem/shell/code-exec. Bad config fails boot, not conversation.

Not yet

  • Real STT/TTS/Opus backends (lazy, need voice-server GPU)
  • Firmware-verified protocol constants (marked # FIRMWARE)
  • Cloudflare tunnel config (external to the bridge)
  • MCP device bridge (Phase 10, after voice MVP)
Description
No description provided
Readme 2.4 MiB
Languages
Python 100%