The llama.cpp endpoint writes the model's hidden reasoning INLINE into
content unless enable_thinking=true is sent (which relocates it to the
separate reasoning_content field). The app sent the field absent, so:
- complete_json got corrupted JSON -> live 500 'Unterminated string char 143'
- persona JSON generation failed (same root cause, earlier incident)
- persona replies were bloated with the reasoning chain
Send the flag on all three call paths (complete/complete_json/
complete_conversation). thinking param retained for API compat, now a
no-op on this provider (always separated — the only correct mode here).
- tests/test_llm_enable_thinking.py pins the flag on all three paths
- full suite: 543 passed; independent reviewer PASS (22 focused passed,
single OpenAI client in the app, both fixed sites confirmed, no stale
tests)
- cleanup stale comments that encoded the backwards understanding
(endpoint 'ignores enable_thinking' — it honors it; that misconception
is what caused this incident)
The Qwen3.8 GGUF reasons internally before answering and the llama.cpp
endpoint ignores enable_thinking — hidden reasoning consumed the whole
400-token persona_reply budget (HTTP 200, content='') -> LLMError -> live
500 'LLM service unavailable'. The 800-token judge JSON had the same
starve and silently degraded to 'pending', so customers who decided to
buy never closed the session.
- persona_reply 400 -> 1200, evaluate_turn 800 -> 1600
- tests/test_llm_budgets.py pins the floors so a regression to the old
values fails loudly
- full suite: 540 passed; independent review PASS (truncation-safe:
partial protocol JSON is rejected, never rendered)
- social trial: silent/gone follow-up outcomes with localized notice
variants; gone is always a loss with possibility-framed debrief
- gone marker persists across judge outages; retries stay on the
deterministic loss path (no roleplay reopen)
- bounded analysis_progress persisted during persona generation,
exposed via safe serializer, polled + rendered as accessible UI
- LLM HTTP timeout 105s -> 180s for Qwen 3.8 (enable_thinking off)
- demo-account test clock made relative to current UTC
- regression coverage: follow-up branches, judge-failure retry,
progress callback, timeout contract, progress serializer
- independent review finding verified stale against final tree
Tests: 532 backend, 29 frontend, prod build green
Feature-word 'persona'/'personas' now reads 'Persona'/'Personas' (capital P)
across the current review checkpoint docs, matching the UI terminology. Code
identifiers, API paths, and schema fields stay lowercase.
Records written before the visibility field existed carry none; the new
fail-closed authorization treated missing visibility as invalid, making every
legacy group unlistable and unreadable. Add resolved_visibility(group) that
derives effective visibility for legacy records only (owner present => private,
absent => public), leaves explicit-malformed visibility fail-closed (None), and
never derives demo/hidden. Apply it at every list, authorization, chat, and
analytics boundary while keeping demo and hidden-preview paths raw and
owner_user_id-based private isolation intact. No persisted data is rewritten.
Backend full suite passes 517; frontend 26/26; production build passes.
Public social signup into OAUTH_DEFAULT_ORG (role user, seat-checked);
email-match links existing active user instead of duplicating. Server-side
provider token validation via stdlib urllib only (no new dep): Google
tokeninfo (aud + email_verified) and Facebook app/debug-token/me (is_valid,
app_id, me.id==user_id). Fail-closed when creds unconfigured, rate-limited
per-IP + per-email, /oauth/config leaks no secrets. Frontend: login buttons
(only enabled providers), GSI + FB SDK on-demand, monochrome glyphs, TH/EN.
Login page shows social buttons only when backend reports provider enabled.
348 backend tests pass (337 + 11 new OAuth), frontend build + 4/4 unit
clean, manual security review PASS. Not pushed (push auto-deploys).
Re-verified staged increment from a clean requirements.lock.txt venv:
- 330 backend tests pass (17/17 in new error_handlers + json_import tests)
- compileall + frontend npm build clean
- git diff --check clean; no secrets in diff
- importer CLI dry-run bootstrap works
Includes JSON HTTPException handler under /api/* and parse-safe static 404
via abort. JSON stores remain runtime-authoritative; production operation
still gated behind operator approval.
Phase 1 (tenant isolation):
- g.org_id set on require_auth; assert_tenant()/current_org_id() choke-point helpers.
- Multi-org provisioning: POST /api/admin/users {new_org:true} (super_admin) creates a
new org + its first admin; GET /api/admin/orgs (super_admin sees all, admin own).
- Fixed latent create_org double-id bug (dict id != store key).
- test_saas_tenant.py: org2 admin blocked from org1 group (403), can't list org1
groups/users, sees only own org; super_admin sees all.
Phase 2 (hardening):
- Rate limit login (per-IP + per-username) + chat send (per-user) to protect LLM cost
and slow brute force; services/rate_limit.py (in-memory + disk, no external deps).
- Audit log data/audit/audit.jsonl on org.create, user.promote_super_admin, analytics.export.
- CSV export now org-scoped (admin exports only own org).
All 8 backend suites pass.