docs: handoff + engineering log (Phase 1 complete)

- docs/HANDOFF.md: resume steps for the 5060 Ti voice-server
- docs/engineering-log.md + dated entry, test-evidence
- git diff --check clean
This commit is contained in:
Macky
2026-10-03 11:57:28 +07:00
parent b8950b38a4
commit 52d9ecfb70
4 changed files with 106 additions and 0 deletions

55
docs/HANDOFF.md Normal file
View File

@@ -0,0 +1,55 @@
# HANDOFF — hermes-xiaozhi-bridge
**Branch:** `main` · **Repo:** `git.moreminimore.com/kunthawat/esp32-server`
**Phase:** 1 complete (testable skeleton, 18 tests green)
**Target next host:** voice-server with RTX 5060 Ti 16GB + V100 32GB
## What exists and works
- `app/main.py` — FastAPI app: `/health` + `/ws/xiaozhi` WebSocket endpoint
- `app/xiaozhi/` — hello/hello-reply protocol, message frames, `Session` state machine, per-connection `DeviceLoop`
- `app/audio/` — `PCMBuffer` (per-utterance), `EnergyVAD` (RMS), Opus: `PassthroughOpus` (test) + `RealOpus` (lazy `import opus`)
- `app/stt/` — `STTEngine` ABC + `MockSTT` + `TyphoonSTT` + `FasterWhisperSTT` (lazy)
- `app/tts/` — `TTSEngine` ABC + `MockTTS` + `JaiTTS` (lazy) + `chunk_text()` (Thai, 60-char hard cap, keeps punctuation) + `TTSQueue` (barge-in cancel)
- `app/hermes/` — `HermesClient` (mock transport now; `openai_http` streaming path ready for Qwen) + `VoiceProfile` guard: asserts `reasoning=="none"` and `tools==["session_search"]` at startup (rejects `terminal`/file tools)
- `app/security/` — device allowlist, SHA-256 token hash
- `app/metrics/` — `LatencyLogger`: speech_start → speech_end → stt_final → hermes_request → qwen_first_token → first_audio_packet → playback_start; `eos_to_first_audio()`
- `app/gpu/` — `GpuMonitor` (VRAM) + `GpuResourceManager.plan()` → FULL/LITE/OFF (pure decision, no process spawning — `voice-full`/`voice-lite`/`voice-off` commands still to wire per plan Phase 13)
**Every backend defaults to `mock` in `config.yaml`** — the bridge boots and passes the full test suite on a CPU-only box. Real backends activate only on the voice-server via config; no code change needed.
## Verified (this host, 2026-10-03)
```bash
.venv/bin/python -m pytest -q # → 18 passed
.venv/bin/python -c "import app.main" # → OK
```
CON-002 checked: no torch/typhoon/whisper/jait/opus/numpy in `sys.modules` after import.
## Do this first on the voice-server (ordered)
1. **Python 3.11–3.12 venv** (this host ran 3.14; prefer 3.12 for CUDA stack compatibility), `pip install -r requirements.txt`
2. **Protocol verification (CON-004) — do this before anything else.** All protocol constants in `config.yaml` under `protocol:` and in `app/xiaozhi/protocol.py` are marked `# FIRMWARE` PoC assumptions:
- `hello_version: 1`, `opus_sample_rate: 16000`, `opus_channels: 1`, `frame_ms: 20`
- Inspect the actual Xiaozhi firmware/OTA used by the ESP32 (its `xai`/`websocket` client) for the real hello schema, auth field name, and Opus framing. `tests/test_core.py::test_hello_parses_firmware_aliases` shows both camelCase and snake_case are accepted — keep the alias layer after verification.
3. **STT**: `stt.engine: typhoon` — point `typhoon_endpoint` at the Typhoon ASR realtime server on the 5060 Ti. Benchmark Thai-English code-switching on the term list in the plan (Qwen, CUDA, vLLM, ComfyUI, …).
4. **TTS**: `tts.engine: jaitts` — point `jaitts_endpoint` at JaiTTS on the 5060 Ti. First-audio latency is the headline metric.
5. **Hermes transport**: `hermes.transport: openai_http` + `base_url`/`api_key` for the Qwen 3.8 vLLM endpoint (V100). Streaming SSE path already implemented in `app/hermes/__init__.py`.
6. **Device auth**: generate a real token per device, `token_hash: sha256(token)` hex in `config.yaml`.
7. **Bind**: keep `host: 127.0.0.1` — the Cloudflare tunnel is the only external surface (REQ-012). Add the tunnel hostname → `127.0.0.1:8765` on the existing tunnel.
## Known sharp edges
- **Chunker**: splits on `.!?。!?ฯ` only, keeps the delimiter with its fragment, hard-slices >60 chars, merges ≤4-char micro-fragments. Do NOT add Thai vowel/tone marks to the split class — it breaks words (bug found in review, see engineering-log).
- **History ownership**: `Session.run_turn` is the single place conversation history is pushed. Don't push from transports.
- **VAD is binary per-frame** (`EnergyVAD.is_speech`); end-of-speech silence-gap logic is the next real work item on the voice-server (PoC uses one utterance = one websocket binary message in tests).
- `openai_http` transport is implemented but **untested live** — mock transport is the tested path.
- uvicorn is in requirements (was not installed on the authoring Mac; it is in `requirements.txt` now).
## Out of scope for Phase 1 (per plan)
MCP device bridge (Phase 10), benchmark suite (Phase 15), GPU command scripts (Phase 13), auto resource manager (deferred — manual `voice-full/lite/off` first).
## Files to read next
`project.md` (requirements) → `plan.md` (status/evidence) → `docs/engineering-log/` (dated decisions) → `app/config.py` (every knob).

16
docs/engineering-log.md Normal file
View File

@@ -0,0 +1,16 @@
# Engineering Log — hermes-xiaozhi-bridge
| Milestone | Status | Last verified | Evidence | Next action |
|---|---|---|---|---|
| Phase 1 skeleton (endpoint, mocks, guards) | complete | 2026-10-03 | `pytest -q` → 18 passed | protocol verification on voice-server |
| STT real (Typhoon) | not started | — | — | wire `typhoon_endpoint` |
| TTS real (JaiTTS) | not started | — | — | wire `jaitts_endpoint` |
| Hermes live (Qwen vLLM) | not started | — | — | `openai_http` transport untested live |
| Firmware protocol lock (CON-004) | in progress | — | `# FIRMWARE` constants are PoC | inspect real firmware |
| EoS→First-Audio benchmark | not started | — | — | Phase 14 logger stages exist |
Guardrails: no secrets in repo · `config.yaml` is the only source of model/token/protocol values (MUST-NOT-001) · `import app` must stay GPU/audio-lib free (CON-002).
## Entry index
- `2026-10-03-phase1-skeleton.md` — Phase 1 complete, 18 tests green, first push

View File

@@ -0,0 +1,32 @@
# 2026-10-03 — Phase 1 skeleton complete (first push)
## Plan status
Phase 1 = complete (testable skeleton). Phase 2+ pending on voice-server.
## What was built
Full repo structure per plan: `app/` (main, config, xiaozhi/, audio/, stt/, tts/, hermes/, security/, metrics/, gpu/), `tests/`, `config.yaml`, `requirements.txt`, `README.md`, `docs/`.
## Key decisions
- **All backends `mock` by default** in `config.yaml` → bridge boots and tests green on CPU-only Mac; real backends lazy-imported, activated by config on voice-server (CON-002 verified: no GPU/audio libs in `sys.modules` at import).
- **Voice profile guard** (`app/hermes/__init__.py::VoiceProfile`) asserts `reasoning == "none"` and `tools == ["session_search"]` at startup → bad config fails boot, not conversation (AC-004).
- **History ownership**: `Session.run_turn` is the single push site for conversation history (a double-push bug found+fixed during test authoring).
- **Chunker** (`app/tts/__init__.py::chunk_text`): split class is `[.!?。!?\u0e3f]` ONLY — Thai vowel/tone marks (U+0E40/41/48) are NOT split points (initial implementation split on them and would have broken every word; caught in review). Delimiter stays attached to its fragment for TTS intonation. 60-char hard cap; ≤4-char micro-fragments merged.
## Evidence
- `pytest -q` → **18 passed** (0.13s) — protocol, config, auth, chunker, barge-in queue, latency, voice-profile guard, GPU planner, VAD, e2e mock spoken turn, app import + live WS handshake+turn.
- `python -c "import app.main"` → OK.
- `chunk_text("สวัสดี. แล้วไง? ครับ")` → `["สวัสดี.", "แล้วไง?", "ครับ"]` (direct call).
- `sys.modules` scan after import → no torch/typhoon/whisper/jait/opus/numpy (CON-002).
- See `docs/test-evidence/2026-10-03-phase1-pytest.txt`.
## Risks / open
- `# FIRMWARE` protocol constants (hello_version=1, frame_ms=20, Opus 16 kHz mono) are PoC assumptions — must verify against real Xiaozhi firmware before Phase 2 (CON-004).
- `openai_http` Hermes transport implemented but untested live (mock is the tested path).
- VAD end-of-speech silence-gap logic is PoC (one utterance = one binary message in tests).
## Next action (voice-server)
1. venv + `pip install -r requirements.txt`
2. Protocol verification against real firmware (CON-004) — first and blocking
3. Wire `typhoon_endpoint`, `jaitts_endpoint`, Hermes `openai_http` base_url/api_key
4. Real device tokens in `config.yaml`
5. Cloudflare tunnel hostname → 127.0.0.1:8765

View File

@@ -0,0 +1,3 @@
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
18 passed, 1 warning in 0.20s