- main() previously used Config() defaults, silently ignoring config.yaml (port + token); now loads config.yaml from repo root (BRIDGE_CONFIG override) - config.yaml: port 8766 (8765 taken on this host), real sha256 token for xiaozhi-main - smoke_live.py: live WS verification (health, auth rejection x2, handshake, full turn) - README: correct run command (python -m app.main) - HANDOFF/plan: Phase 1.5 live-verified evidence
5.6 KiB
HANDOFF — hermes-xiaozhi-bridge
Branch: main · Repo: git.moreminimore.com/kunthawat/esp32-server
Phase: 1 complete (testable skeleton, 18 tests green)
Target next host: voice-server with RTX 5060 Ti 16GB + V100 32GB
What exists and works
app/main.py— FastAPI app:/health+/ws/xiaozhiWebSocket endpointapp/xiaozhi/— hello/hello-reply protocol, message frames,Sessionstate machine, per-connectionDeviceLoopapp/audio/—PCMBuffer(per-utterance),EnergyVAD(RMS), Opus:PassthroughOpus(test) +RealOpus(lazyimport opus)app/stt/—STTEngineABC +MockSTT+TyphoonSTT+FasterWhisperSTT(lazy)app/tts/—TTSEngineABC +MockTTS+JaiTTS(lazy) +chunk_text()(Thai, 60-char hard cap, keeps punctuation) +TTSQueue(barge-in cancel)app/hermes/—HermesClient(mock transport now;openai_httpstreaming path ready for Qwen) +VoiceProfileguard: assertsreasoning=="none"andtools==["session_search"]at startup (rejectsterminal/file tools)app/security/— device allowlist, SHA-256 token hashapp/metrics/—LatencyLogger: speech_start → speech_end → stt_final → hermes_request → qwen_first_token → first_audio_packet → playback_start;eos_to_first_audio()app/gpu/—GpuMonitor(VRAM) +GpuResourceManager.plan()→ FULL/LITE/OFF (pure decision, no process spawning —voice-full/voice-lite/voice-offcommands still to wire per plan Phase 13)
Every backend defaults to mock in config.yaml — the bridge boots and passes the full test suite on a CPU-only box. Real backends activate only on the voice-server via config; no code change needed.
Verified (this host, 2026-10-03)
.venv/bin/python -m pytest -q # → 18 passed
.venv/bin/python -c "import app.main" # → OK
.venv/bin/python -m app.main # → serves /health + /ws/xiaozhi on config.yaml host:port
.venv/bin/python smoke_live.py # → live WS smoke: health, bad-token/bad-device rejection,
# handshake, spoken turn (stt→state→text→audio) — PASS
CON-002 checked: no torch/typhoon/whisper/jait/opus/numpy in sys.modules after import.
Bug found+fixed on this host: main() ignored config.yaml (used Config() defaults —
wrong port + zero-hashes token); now loads config.yaml from the repo root
(override with BRIDGE_CONFIG). Real device token issued for xiaozhi-main.
Do this first on the voice-server (ordered)
- Python 3.11–3.12 venv (this host ran 3.14; prefer 3.12 for CUDA stack compatibility),
pip install -r requirements.txt - Protocol verification (CON-004) — do this before anything else. All protocol constants in
config.yamlunderprotocol:and inapp/xiaozhi/protocol.pyare marked# FIRMWAREPoC assumptions:hello_version: 1,opus_sample_rate: 16000,opus_channels: 1,frame_ms: 20- Inspect the actual Xiaozhi firmware/OTA used by the ESP32 (its
xai/websocketclient) for the real hello schema, auth field name, and Opus framing.tests/test_core.py::test_hello_parses_firmware_aliasesshows both camelCase and snake_case are accepted — keep the alias layer after verification.
- STT:
stt.engine: typhoon— pointtyphoon_endpointat the Typhoon ASR realtime server on the 5060 Ti. Benchmark Thai-English code-switching on the term list in the plan (Qwen, CUDA, vLLM, ComfyUI, …). - TTS:
tts.engine: jaitts— pointjaitts_endpointat JaiTTS on the 5060 Ti. First-audio latency is the headline metric. - Hermes transport:
hermes.transport: openai_http+base_url/api_keyfor the Qwen 3.8 vLLM endpoint (V100). Streaming SSE path already implemented inapp/hermes/__init__.py. - Device auth: generate a real token per device,
token_hash: sha256(token)hex inconfig.yaml.- Done on this host (2026-10-03):
xiaozhi-maintoken =8a07643d106f9282a8f2eb1161da4f73657ce99b6602b98df89d011b8caf48a2(store in the ESP32 firmware config;token_hashin config.yaml is its SHA-256). - ⚠️ config.yaml is committed to the repo — this token is public. Rotate it (new token, new hash) before the device leaves the lab, or move config to a gitignored overlay.
- Done on this host (2026-10-03):
- Bind: keep
host: 127.0.0.1— the Cloudflare tunnel is the only external surface (REQ-012). Point the tunnel hostname at127.0.0.1:<port>where<port>is the value inconfig.yaml(this host uses 8766 — 8765 is taken by another local service).
Known sharp edges
- Chunker: splits on
.!?。!?ฯonly, keeps the delimiter with its fragment, hard-slices >60 chars, merges ≤4-char micro-fragments. Do NOT add Thai vowel/tone marks to the split class — it breaks words (bug found in review, see engineering-log). - History ownership:
Session.run_turnis the single place conversation history is pushed. Don't push from transports. - VAD is binary per-frame (
EnergyVAD.is_speech); end-of-speech silence-gap logic is the next real work item on the voice-server (PoC uses one utterance = one websocket binary message in tests). openai_httptransport is implemented but untested live — mock transport is the tested path.- uvicorn is in requirements (was not installed on the authoring Mac; it is in
requirements.txtnow).
Out of scope for Phase 1 (per plan)
MCP device bridge (Phase 10), benchmark suite (Phase 15), GPU command scripts (Phase 13), auto resource manager (deferred — manual voice-full/lite/off first).
Files to read next
project.md (requirements) → plan.md (status/evidence) → docs/engineering-log/ (dated decisions) → app/config.py (every knob).