Files
esp32-server/docs/HANDOFF.md
Macky 52d9ecfb70 docs: handoff + engineering log (Phase 1 complete)
- docs/HANDOFF.md: resume steps for the 5060 Ti voice-server
- docs/engineering-log.md + dated entry, test-evidence
- git diff --check clean
2026-10-03 11:57:28 +07:00

4.6 KiB
Raw Blame History

HANDOFF — hermes-xiaozhi-bridge

Branch: main · Repo: git.moreminimore.com/kunthawat/esp32-server Phase: 1 complete (testable skeleton, 18 tests green) Target next host: voice-server with RTX 5060 Ti 16GB + V100 32GB

What exists and works

  • app/main.py — FastAPI app: /health + /ws/xiaozhi WebSocket endpoint
  • app/xiaozhi/ — hello/hello-reply protocol, message frames, Session state machine, per-connection DeviceLoop
  • app/audio/ — PCMBuffer (per-utterance), EnergyVAD (RMS), Opus: PassthroughOpus (test) + RealOpus (lazy import opus)
  • app/stt/ — STTEngine ABC + MockSTT + TyphoonSTT + FasterWhisperSTT (lazy)
  • app/tts/ — TTSEngine ABC + MockTTS + JaiTTS (lazy) + chunk_text() (Thai, 60-char hard cap, keeps punctuation) + TTSQueue (barge-in cancel)
  • app/hermes/ — HermesClient (mock transport now; openai_http streaming path ready for Qwen) + VoiceProfile guard: asserts reasoning=="none" and tools==["session_search"] at startup (rejects terminal/file tools)
  • app/security/ — device allowlist, SHA-256 token hash
  • app/metrics/ — LatencyLogger: speech_start → speech_end → stt_final → hermes_request → qwen_first_token → first_audio_packet → playback_start; eos_to_first_audio()
  • app/gpu/ — GpuMonitor (VRAM) + GpuResourceManager.plan() → FULL/LITE/OFF (pure decision, no process spawning — voice-full/voice-lite/voice-off commands still to wire per plan Phase 13)

Every backend defaults to mock in config.yaml — the bridge boots and passes the full test suite on a CPU-only box. Real backends activate only on the voice-server via config; no code change needed.

Verified (this host, 2026-10-03)

.venv/bin/python -m pytest -q          # → 18 passed
.venv/bin/python -c "import app.main"  # → OK

CON-002 checked: no torch/typhoon/whisper/jait/opus/numpy in sys.modules after import.

Do this first on the voice-server (ordered)

  1. Python 3.11–3.12 venv (this host ran 3.14; prefer 3.12 for CUDA stack compatibility), pip install -r requirements.txt
  2. Protocol verification (CON-004) — do this before anything else. All protocol constants in config.yaml under protocol: and in app/xiaozhi/protocol.py are marked # FIRMWARE PoC assumptions:
    • hello_version: 1, opus_sample_rate: 16000, opus_channels: 1, frame_ms: 20
    • Inspect the actual Xiaozhi firmware/OTA used by the ESP32 (its xai/websocket client) for the real hello schema, auth field name, and Opus framing. tests/test_core.py::test_hello_parses_firmware_aliases shows both camelCase and snake_case are accepted — keep the alias layer after verification.
  3. STT: stt.engine: typhoon — point typhoon_endpoint at the Typhoon ASR realtime server on the 5060 Ti. Benchmark Thai-English code-switching on the term list in the plan (Qwen, CUDA, vLLM, ComfyUI, …).
  4. TTS: tts.engine: jaitts — point jaitts_endpoint at JaiTTS on the 5060 Ti. First-audio latency is the headline metric.
  5. Hermes transport: hermes.transport: openai_http + base_url/api_key for the Qwen 3.8 vLLM endpoint (V100). Streaming SSE path already implemented in app/hermes/__init__.py.
  6. Device auth: generate a real token per device, token_hash: sha256(token) hex in config.yaml.
  7. Bind: keep host: 127.0.0.1 — the Cloudflare tunnel is the only external surface (REQ-012). Add the tunnel hostname → 127.0.0.1:8765 on the existing tunnel.

Known sharp edges

  • Chunker: splits on .!?。!?ฯ only, keeps the delimiter with its fragment, hard-slices >60 chars, merges ≤4-char micro-fragments. Do NOT add Thai vowel/tone marks to the split class — it breaks words (bug found in review, see engineering-log).
  • History ownership: Session.run_turn is the single place conversation history is pushed. Don't push from transports.
  • VAD is binary per-frame (EnergyVAD.is_speech); end-of-speech silence-gap logic is the next real work item on the voice-server (PoC uses one utterance = one websocket binary message in tests).
  • openai_http transport is implemented but untested live — mock transport is the tested path.
  • uvicorn is in requirements (was not installed on the authoring Mac; it is in requirements.txt now).

Out of scope for Phase 1 (per plan)

MCP device bridge (Phase 10), benchmark suite (Phase 15), GPU command scripts (Phase 13), auto resource manager (deferred — manual voice-full/lite/off first).

project.md (requirements) → plan.md (status/evidence) → docs/engineering-log/ (dated decisions) → app/config.py (every knob).