- C1: stop feeding the unconsumed barge-in TTSQueue in the reply path (deadlocked after 64 chunks on a real long reply)
- C2: encode reply Opus at the TTS/codec sample rate (24 kHz), not hardcoded 16 kHz; encoder now rate-generic
- M3: bound the mic PCM buffer to the 60 s max utterance
- M4: tts_sentence_start per speakable sentence, not per LLM token
- Minor: send tts.stop on a failed turn; constant-time token compare
- Sec: untrack + gitignore config.yaml; add config.example.yaml with placeholder secrets
- Tests: 23 pass (2 new regressions: long-reply deadlock, rate-generic encoder)
Device streams continuous 20ms Opus frames with no turn markers. The loop
now decodes Opus, feeds the VAD, fires one turn on end-of-speech, and
replies with the firmware JSON contract (tts.state / stt.text) + Opus audio.
- app/audio/opus.py: RealOpus/Opus16kEncoder/PassthroughOpus; fix add_dll_directory
to use absolute paths (WinError 87)
- app/xiaozhi/websocket.py: DeviceLoop.run() = decode -> VAD -> run_turn -> reply
- app/config.py: AudioConfig (opus kind + VAD thresholds)
- config.yaml: audio.opus=real, VAD 5/8/3000 frames
- tests/test_app.py: utterance = loud burst + silence; asserts stt + audio + tts stop
- requirements.txt: opuslib>=3
21/21 tests; live handshake verified via zhi.moreminimore.com on venv python.
- main() previously used Config() defaults, silently ignoring config.yaml
(port + token); now loads config.yaml from repo root (BRIDGE_CONFIG override)
- config.yaml: port 8766 (8765 taken on this host), real sha256 token for xiaozhi-main
- smoke_live.py: live WS verification (health, auth rejection x2, handshake, full turn)
- README: correct run command (python -m app.main)
- HANDOFF/plan: Phase 1.5 live-verified evidence