10 Commits

Author SHA1 Message Date
e96498b67e fix: resolve fresh-eyes review criticals + majors on the audio path
- C1: stop feeding the unconsumed barge-in TTSQueue in the reply path (deadlocked after 64 chunks on a real long reply)
- C2: encode reply Opus at the TTS/codec sample rate (24 kHz), not hardcoded 16 kHz; encoder now rate-generic
- M3: bound the mic PCM buffer to the 60 s max utterance
- M4: tts_sentence_start per speakable sentence, not per LLM token
- Minor: send tts.stop on a failed turn; constant-time token compare
- Sec: untrack + gitignore config.yaml; add config.example.yaml with placeholder secrets
- Tests: 23 pass (2 new regressions: long-reply deadlock, rate-generic encoder)
2026-10-03 21:55:32 +07:00
261b0f3e91 feat: wire Opus decode + VAD turn detection to firmware-compatible reply
Device streams continuous 20ms Opus frames with no turn markers. The loop
now decodes Opus, feeds the VAD, fires one turn on end-of-speech, and
replies with the firmware JSON contract (tts.state / stt.text) + Opus audio.

- app/audio/opus.py: RealOpus/Opus16kEncoder/PassthroughOpus; fix add_dll_directory
  to use absolute paths (WinError 87)
- app/xiaozhi/websocket.py: DeviceLoop.run() = decode -> VAD -> run_turn -> reply
- app/config.py: AudioConfig (opus kind + VAD thresholds)
- config.yaml: audio.opus=real, VAD 5/8/3000 frames
- tests/test_app.py: utterance = loud burst + silence; asserts stt + audio + tts stop
- requirements.txt: opuslib>=3

21/21 tests; live handshake verified via zhi.moreminimore.com on venv python.
2026-10-03 20:35:38 +07:00
3bfeff29bd docs: Phase 2 — firmware handshake/OTA spec verified from source, build in progress 2026-10-03 16:36:41 +07:00
a582c1ff36 feat: firmware-compatible handshake (header auth, transport=websocket) + /ota onboarding endpoint
- auth from Authorization/Device-Id headers (firmware v2.5.0), frame fallback kept
- server hello carries transport=websocket + session_id (firmware ParseServerHello)
- /ota GET+POST returns {websocket:{url,token,version}, server_time:{timestamp,timezone_offset}, firmware:{version}}
- config.yaml: device registered by MAC, ota block added
- 21 tests pass (added OTA + header-auth coverage)
2026-10-03 16:35:41 +07:00
14394859aa docs: CON-003 override — custom ESP-IDF build (OTA server + Thai) 2026-10-03 15:50:26 +07:00
662c663cda docs: phase 1.6 — live device path verified, push infra fixed 2026-10-03 15:21:22 +07:00
e91d185996 fix: module-level app loads config.yaml (uvicorn app.main:app); bind 0.0.0.0 for LAN/device reachability 2026-10-03 15:19:09 +07:00
b00312cbe5 fix: load config.yaml on boot; issue real device token; live smoke test
- main() previously used Config() defaults, silently ignoring config.yaml
  (port + token); now loads config.yaml from repo root (BRIDGE_CONFIG override)
- config.yaml: port 8766 (8765 taken on this host), real sha256 token for xiaozhi-main
- smoke_live.py: live WS verification (health, auth rejection x2, handshake, full turn)
- README: correct run command (python -m app.main)
- HANDOFF/plan: Phase 1.5 live-verified evidence
2026-10-03 12:12:29 +07:00
Macky
52d9ecfb70 docs: handoff + engineering log (Phase 1 complete)
- docs/HANDOFF.md: resume steps for the 5060 Ti voice-server
- docs/engineering-log.md + dated entry, test-evidence
- git diff --check clean
2026-10-03 11:57:28 +07:00
Macky
b8950b38a4 feat: Phase 1 testable bridge skeleton (18 tests green)
Xiaozhi WebSocket endpoint with handshake/auth, mock STT/TTS/Hermes
backends, Thai chunker, barge-in queue, latency logger, GPU planner,
voice-profile guard (reasoning=none, session_search only).

CON-002: no GPU/audio libs loaded at import. MUST-NOT-001: all model
names/tokens from config.
2026-10-03 11:45:31 +07:00