- complete_json(): log model/max_tokens + raw provider text (1500 cap) when json.loads fails, so a truncated judge reply is diagnosable from the log - complete()/complete_conversation(): log reasoning_content head on empty content (hidden reasoning starves the shared max_tokens cap) - simulator.judge(): max_tokens 2000 -> 6000 (server accepts 8000) - tests: pin judge budget; new test_llm_failure_logging.py pins the logging (autouse fixture restores app.llm logger disabled by alembic fileConfig) 547 passed. Incident #4 (500 char 125).
3.7 KiB
Sales Trainer — Current Execution Plan
Status
Incident #4 fix committed locally (2026-10-02): raw-output logging in llm.py (invalid JSON → raw provider text capped 1500 chars; empty content → reasoning_content head), judge max_tokens 2000→6000, 4 regression tests. 547 passed; independent reviewer REJECT (2 items) resolved and re-verified. Awaiting owner push approval.
Current phase / active task
Incident #4: fix implemented + reviewed + committed locally. Awaiting owner push approval → EasyPanel auto-deploy (~3 min) → live verify.
Last update
2026-10-02. Incident #4 (500 judge JSON truncated, char 125): raw-output logging in llm.py + judge budget 2000→6000 + 4 regression tests. 547 passed. Committed locally, awaiting owner push approval.
Next action
Owner approves push → EasyPanel auto-deploy (~3 min) → live verify: next chat send → 200 + clean judge outcome. If a JSON 500 returns, the new log line carries the raw provider output (capped 1500 chars) and the model/max_tokens context.
Task breakdown
- TASK-001 — Trace follow-up/session state machine, progress/polling, one-shot lifecycle, timeout chain, and historic error path. Covers REQ-001–REQ-007; exact historic error remains unverified because local logs/reproduction are unavailable.
- TASK-002 — Add regression coverage for follow-up branches, notice variants, gone debrief/loss, and judge retry. Covers REQ-001–REQ-003, AC-001.
- TASK-003 — Verify existing completed-persona API/UI lifecycle and one-shot guards; no extra change needed. Covers REQ-004, AC-001.
- TASK-004 — Persist bounded progress, expose safe serializer, poll, and display accessible generation progress. Covers REQ-005, AC-002.
- TASK-005 — Increase LLM timeout to 180s, retain finite 30-minute async polling, investigate error from code evidence. Covers REQ-006–REQ-007, AC-003/AC-005.
- TASK-006 — Full backend/frontend checks and security scan passed; independent reviewer verdict received and verified. Single REJECT finding (gone session reopenable after judge failure) confirmed stale — reviewer read a pre-fix
chat_routes.pysnapshot; current tree has thefollowup_phase == "gone"retry branch (chat_routes.py:888) plus both regression tests, which pass. Full suite re-run green post-review (532). No other confirmed blocker. Covers AC-001–AC-005. - TASK-007 — Engineering log, handoff, test evidence, and this plan updated with the reviewer verdict. Covers AC-004/AC-005.
- TASK-008 — Incident #4 (2026-10-02): log raw LLM output on JSON parse failure + reasoning head on empty content (llm.py); judge
max_tokens2000→6000 (simulator.py); 4 regression tests (test_llm_budgets.py + new test_llm_failure_logging.py). 547 passed.
Blockers
None known. Do not push/deploy without owner approval. Do not access .env or production data.
Dependencies
TASK-002–TASK-005 depend on TASK-001. TASK-006 depends on all implementation tasks; TASK-007 depends on verification results.
Requirement coverage
REQ-001/002/003 → TASK-002; REQ-004 → TASK-003; REQ-005 → TASK-004; REQ-006/007 → TASK-005; all acceptance criteria → TASK-006/007.
Verification and evidence
- RED/GREEN targeted pytest for each backend behavior before implementation.
- Run complete backend pytest suite,
python -m compileall -q app tests, frontend unit tests and production build. - Inspect
git diff --check; scan changed lines for secrets, injection, unsafe execution, and debug leftovers. - Independently review exact current diff with diagnosis reviewer; do not treat timeout/no verdict as approval.
- No production/real-provider claims unless explicitly verified; no push/deploy.