- docs/HANDOFF.md: resume steps for the 5060 Ti voice-server - docs/engineering-log.md + dated entry, test-evidence - git diff --check clean
4.6 KiB
4.6 KiB
HANDOFF — hermes-xiaozhi-bridge
Branch: main · Repo: git.moreminimore.com/kunthawat/esp32-server
Phase: 1 complete (testable skeleton, 18 tests green)
Target next host: voice-server with RTX 5060 Ti 16GB + V100 32GB
What exists and works
app/main.py— FastAPI app:/health+/ws/xiaozhiWebSocket endpointapp/xiaozhi/— hello/hello-reply protocol, message frames,Sessionstate machine, per-connectionDeviceLoopapp/audio/—PCMBuffer(per-utterance),EnergyVAD(RMS), Opus:PassthroughOpus(test) +RealOpus(lazyimport opus)app/stt/—STTEngineABC +MockSTT+TyphoonSTT+FasterWhisperSTT(lazy)app/tts/—TTSEngineABC +MockTTS+JaiTTS(lazy) +chunk_text()(Thai, 60-char hard cap, keeps punctuation) +TTSQueue(barge-in cancel)app/hermes/—HermesClient(mock transport now;openai_httpstreaming path ready for Qwen) +VoiceProfileguard: assertsreasoning=="none"andtools==["session_search"]at startup (rejectsterminal/file tools)app/security/— device allowlist, SHA-256 token hashapp/metrics/—LatencyLogger: speech_start → speech_end → stt_final → hermes_request → qwen_first_token → first_audio_packet → playback_start;eos_to_first_audio()app/gpu/—GpuMonitor(VRAM) +GpuResourceManager.plan()→ FULL/LITE/OFF (pure decision, no process spawning —voice-full/voice-lite/voice-offcommands still to wire per plan Phase 13)
Every backend defaults to mock in config.yaml — the bridge boots and passes the full test suite on a CPU-only box. Real backends activate only on the voice-server via config; no code change needed.
Verified (this host, 2026-10-03)
.venv/bin/python -m pytest -q # → 18 passed
.venv/bin/python -c "import app.main" # → OK
CON-002 checked: no torch/typhoon/whisper/jait/opus/numpy in sys.modules after import.
Do this first on the voice-server (ordered)
- Python 3.11–3.12 venv (this host ran 3.14; prefer 3.12 for CUDA stack compatibility),
pip install -r requirements.txt - Protocol verification (CON-004) — do this before anything else. All protocol constants in
config.yamlunderprotocol:and inapp/xiaozhi/protocol.pyare marked# FIRMWAREPoC assumptions:hello_version: 1,opus_sample_rate: 16000,opus_channels: 1,frame_ms: 20- Inspect the actual Xiaozhi firmware/OTA used by the ESP32 (its
xai/websocketclient) for the real hello schema, auth field name, and Opus framing.tests/test_core.py::test_hello_parses_firmware_aliasesshows both camelCase and snake_case are accepted — keep the alias layer after verification.
- STT:
stt.engine: typhoon— pointtyphoon_endpointat the Typhoon ASR realtime server on the 5060 Ti. Benchmark Thai-English code-switching on the term list in the plan (Qwen, CUDA, vLLM, ComfyUI, …). - TTS:
tts.engine: jaitts— pointjaitts_endpointat JaiTTS on the 5060 Ti. First-audio latency is the headline metric. - Hermes transport:
hermes.transport: openai_http+base_url/api_keyfor the Qwen 3.8 vLLM endpoint (V100). Streaming SSE path already implemented inapp/hermes/__init__.py. - Device auth: generate a real token per device,
token_hash: sha256(token)hex inconfig.yaml. - Bind: keep
host: 127.0.0.1— the Cloudflare tunnel is the only external surface (REQ-012). Add the tunnel hostname →127.0.0.1:8765on the existing tunnel.
Known sharp edges
- Chunker: splits on
.!?。!?ฯonly, keeps the delimiter with its fragment, hard-slices >60 chars, merges ≤4-char micro-fragments. Do NOT add Thai vowel/tone marks to the split class — it breaks words (bug found in review, see engineering-log). - History ownership:
Session.run_turnis the single place conversation history is pushed. Don't push from transports. - VAD is binary per-frame (
EnergyVAD.is_speech); end-of-speech silence-gap logic is the next real work item on the voice-server (PoC uses one utterance = one websocket binary message in tests). openai_httptransport is implemented but untested live — mock transport is the tested path.- uvicorn is in requirements (was not installed on the authoring Mac; it is in
requirements.txtnow).
Out of scope for Phase 1 (per plan)
MCP device bridge (Phase 10), benchmark suite (Phase 15), GPU command scripts (Phase 13), auto resource manager (deferred — manual voice-full/lite/off first).
Files to read next
project.md (requirements) → plan.md (status/evidence) → docs/engineering-log/ (dated decisions) → app/config.py (every knob).