Files
sales-trainer/plan.md
Macky 9df2938557 fix(llm): log raw provider output on JSON failure, raise judge budget to 6000
- complete_json(): log model/max_tokens + raw provider text (1500 cap) when
  json.loads fails, so a truncated judge reply is diagnosable from the log
- complete()/complete_conversation(): log reasoning_content head on empty
  content (hidden reasoning starves the shared max_tokens cap)
- simulator.judge(): max_tokens 2000 -> 6000 (server accepts 8000)
- tests: pin judge budget; new test_llm_failure_logging.py pins the logging
  (autouse fixture restores app.llm logger disabled by alembic fileConfig)

547 passed. Incident #4 (500 char 125).
2026-10-02 22:30:28 +07:00

3.7 KiB
Raw Permalink Blame History

Sales Trainer — Current Execution Plan

Status

Incident #4 fix committed locally (2026-10-02): raw-output logging in llm.py (invalid JSON → raw provider text capped 1500 chars; empty content → reasoning_content head), judge max_tokens 2000→6000, 4 regression tests. 547 passed; independent reviewer REJECT (2 items) resolved and re-verified. Awaiting owner push approval.

Current phase / active task

Incident #4: fix implemented + reviewed + committed locally. Awaiting owner push approval → EasyPanel auto-deploy (~3 min) → live verify.

Last update

2026-10-02. Incident #4 (500 judge JSON truncated, char 125): raw-output logging in llm.py + judge budget 2000→6000 + 4 regression tests. 547 passed. Committed locally, awaiting owner push approval.

Next action

Owner approves push → EasyPanel auto-deploy (~3 min) → live verify: next chat send → 200 + clean judge outcome. If a JSON 500 returns, the new log line carries the raw provider output (capped 1500 chars) and the model/max_tokens context.

Task breakdown

  • TASK-001 — Trace follow-up/session state machine, progress/polling, one-shot lifecycle, timeout chain, and historic error path. Covers REQ-001–REQ-007; exact historic error remains unverified because local logs/reproduction are unavailable.
  • TASK-002 — Add regression coverage for follow-up branches, notice variants, gone debrief/loss, and judge retry. Covers REQ-001–REQ-003, AC-001.
  • TASK-003 — Verify existing completed-persona API/UI lifecycle and one-shot guards; no extra change needed. Covers REQ-004, AC-001.
  • TASK-004 — Persist bounded progress, expose safe serializer, poll, and display accessible generation progress. Covers REQ-005, AC-002.
  • TASK-005 — Increase LLM timeout to 180s, retain finite 30-minute async polling, investigate error from code evidence. Covers REQ-006–REQ-007, AC-003/AC-005.
  • TASK-006 — Full backend/frontend checks and security scan passed; independent reviewer verdict received and verified. Single REJECT finding (gone session reopenable after judge failure) confirmed stale — reviewer read a pre-fix chat_routes.py snapshot; current tree has the followup_phase == "gone" retry branch (chat_routes.py:888) plus both regression tests, which pass. Full suite re-run green post-review (532). No other confirmed blocker. Covers AC-001–AC-005.
  • TASK-007 — Engineering log, handoff, test evidence, and this plan updated with the reviewer verdict. Covers AC-004/AC-005.
  • TASK-008 — Incident #4 (2026-10-02): log raw LLM output on JSON parse failure + reasoning head on empty content (llm.py); judge max_tokens 2000→6000 (simulator.py); 4 regression tests (test_llm_budgets.py + new test_llm_failure_logging.py). 547 passed.

Blockers

None known. Do not push/deploy without owner approval. Do not access .env or production data.

Dependencies

TASK-002–TASK-005 depend on TASK-001. TASK-006 depends on all implementation tasks; TASK-007 depends on verification results.

Requirement coverage

REQ-001/002/003 → TASK-002; REQ-004 → TASK-003; REQ-005 → TASK-004; REQ-006/007 → TASK-005; all acceptance criteria → TASK-006/007.

Verification and evidence

  • RED/GREEN targeted pytest for each backend behavior before implementation.
  • Run complete backend pytest suite, python -m compileall -q app tests, frontend unit tests and production build.
  • Inspect git diff --check; scan changed lines for secrets, injection, unsafe execution, and debug leftovers.
  • Independently review exact current diff with diagnosis reviewer; do not treat timeout/no verdict as approval.
  • No production/real-provider claims unless explicitly verified; no push/deploy.