Files
esp32-server/docs/HANDOFF.md
kunthawat b00312cbe5 fix: load config.yaml on boot; issue real device token; live smoke test
- main() previously used Config() defaults, silently ignoring config.yaml
  (port + token); now loads config.yaml from repo root (BRIDGE_CONFIG override)
- config.yaml: port 8766 (8765 taken on this host), real sha256 token for xiaozhi-main
- smoke_live.py: live WS verification (health, auth rejection x2, handshake, full turn)
- README: correct run command (python -m app.main)
- HANDOFF/plan: Phase 1.5 live-verified evidence
2026-10-03 12:12:29 +07:00

5.6 KiB
Raw Blame History

HANDOFF — hermes-xiaozhi-bridge

Branch: main · Repo: git.moreminimore.com/kunthawat/esp32-server Phase: 1 complete (testable skeleton, 18 tests green) Target next host: voice-server with RTX 5060 Ti 16GB + V100 32GB

What exists and works

  • app/main.py — FastAPI app: /health + /ws/xiaozhi WebSocket endpoint
  • app/xiaozhi/ — hello/hello-reply protocol, message frames, Session state machine, per-connection DeviceLoop
  • app/audio/ — PCMBuffer (per-utterance), EnergyVAD (RMS), Opus: PassthroughOpus (test) + RealOpus (lazy import opus)
  • app/stt/ — STTEngine ABC + MockSTT + TyphoonSTT + FasterWhisperSTT (lazy)
  • app/tts/ — TTSEngine ABC + MockTTS + JaiTTS (lazy) + chunk_text() (Thai, 60-char hard cap, keeps punctuation) + TTSQueue (barge-in cancel)
  • app/hermes/ — HermesClient (mock transport now; openai_http streaming path ready for Qwen) + VoiceProfile guard: asserts reasoning=="none" and tools==["session_search"] at startup (rejects terminal/file tools)
  • app/security/ — device allowlist, SHA-256 token hash
  • app/metrics/ — LatencyLogger: speech_start → speech_end → stt_final → hermes_request → qwen_first_token → first_audio_packet → playback_start; eos_to_first_audio()
  • app/gpu/ — GpuMonitor (VRAM) + GpuResourceManager.plan() → FULL/LITE/OFF (pure decision, no process spawning — voice-full/voice-lite/voice-off commands still to wire per plan Phase 13)

Every backend defaults to mock in config.yaml — the bridge boots and passes the full test suite on a CPU-only box. Real backends activate only on the voice-server via config; no code change needed.

Verified (this host, 2026-10-03)

.venv/bin/python -m pytest -q          # → 18 passed
.venv/bin/python -c "import app.main"  # → OK
.venv/bin/python -m app.main           # → serves /health + /ws/xiaozhi on config.yaml host:port
.venv/bin/python smoke_live.py         # → live WS smoke: health, bad-token/bad-device rejection,
                                       #   handshake, spoken turn (stt→state→text→audio) — PASS

CON-002 checked: no torch/typhoon/whisper/jait/opus/numpy in sys.modules after import. Bug found+fixed on this host: main() ignored config.yaml (used Config() defaults — wrong port + zero-hashes token); now loads config.yaml from the repo root (override with BRIDGE_CONFIG). Real device token issued for xiaozhi-main.

Do this first on the voice-server (ordered)

  1. Python 3.11–3.12 venv (this host ran 3.14; prefer 3.12 for CUDA stack compatibility), pip install -r requirements.txt
  2. Protocol verification (CON-004) — do this before anything else. All protocol constants in config.yaml under protocol: and in app/xiaozhi/protocol.py are marked # FIRMWARE PoC assumptions:
    • hello_version: 1, opus_sample_rate: 16000, opus_channels: 1, frame_ms: 20
    • Inspect the actual Xiaozhi firmware/OTA used by the ESP32 (its xai/websocket client) for the real hello schema, auth field name, and Opus framing. tests/test_core.py::test_hello_parses_firmware_aliases shows both camelCase and snake_case are accepted — keep the alias layer after verification.
  3. STT: stt.engine: typhoon — point typhoon_endpoint at the Typhoon ASR realtime server on the 5060 Ti. Benchmark Thai-English code-switching on the term list in the plan (Qwen, CUDA, vLLM, ComfyUI, …).
  4. TTS: tts.engine: jaitts — point jaitts_endpoint at JaiTTS on the 5060 Ti. First-audio latency is the headline metric.
  5. Hermes transport: hermes.transport: openai_http + base_url/api_key for the Qwen 3.8 vLLM endpoint (V100). Streaming SSE path already implemented in app/hermes/__init__.py.
  6. Device auth: generate a real token per device, token_hash: sha256(token) hex in config.yaml.
    • Done on this host (2026-10-03): xiaozhi-main token = 8a07643d106f9282a8f2eb1161da4f73657ce99b6602b98df89d011b8caf48a2 (store in the ESP32 firmware config; token_hash in config.yaml is its SHA-256).
    • ⚠️ config.yaml is committed to the repo — this token is public. Rotate it (new token, new hash) before the device leaves the lab, or move config to a gitignored overlay.
  7. Bind: keep host: 127.0.0.1 — the Cloudflare tunnel is the only external surface (REQ-012). Point the tunnel hostname at 127.0.0.1:<port> where <port> is the value in config.yaml (this host uses 8766 — 8765 is taken by another local service).

Known sharp edges

  • Chunker: splits on .!?。!?ฯ only, keeps the delimiter with its fragment, hard-slices >60 chars, merges ≤4-char micro-fragments. Do NOT add Thai vowel/tone marks to the split class — it breaks words (bug found in review, see engineering-log).
  • History ownership: Session.run_turn is the single place conversation history is pushed. Don't push from transports.
  • VAD is binary per-frame (EnergyVAD.is_speech); end-of-speech silence-gap logic is the next real work item on the voice-server (PoC uses one utterance = one websocket binary message in tests).
  • openai_http transport is implemented but untested live — mock transport is the tested path.
  • uvicorn is in requirements (was not installed on the authoring Mac; it is in requirements.txt now).

Out of scope for Phase 1 (per plan)

MCP device bridge (Phase 10), benchmark suite (Phase 15), GPU command scripts (Phase 13), auto resource manager (deferred — manual voice-full/lite/off first).

project.md (requirements) → plan.md (status/evidence) → docs/engineering-log/ (dated decisions) → app/config.py (every knob).