The llama.cpp endpoint writes the model's hidden reasoning INLINE into
content unless enable_thinking=true is sent (which relocates it to the
separate reasoning_content field). The app sent the field absent, so:
- complete_json got corrupted JSON -> live 500 'Unterminated string char 143'
- persona JSON generation failed (same root cause, earlier incident)
- persona replies were bloated with the reasoning chain
Send the flag on all three call paths (complete/complete_json/
complete_conversation). thinking param retained for API compat, now a
no-op on this provider (always separated — the only correct mode here).
- tests/test_llm_enable_thinking.py pins the flag on all three paths
- full suite: 543 passed; independent reviewer PASS (22 focused passed,
single OpenAI client in the app, both fixed sites confirmed, no stale
tests)
- cleanup stale comments that encoded the backwards understanding
(endpoint 'ignores enable_thinking' — it honors it; that misconception
is what caused this incident)