The paper presents a model-agnostic governance architecture for multi-turn mental health support. It combines contextual risk detection, reasoning-based verification, and protocol-guided response generation so that safety behavior can adapt as conversational risk evolves. On synthetic conversations grounded in real-world mental health narratives, evaluations with GPT-5-chat and Qwen3.5-27B reported 0.92 sensitivity and 0.85 specificity. Clinician-preferred escalation responses increased by 25.6–59.2 percentage points while rapport and connection were preserved. The authors also report stable performance across conversation lengths and generalization across proprietary and open-source models.
No heat snapshots are available in the last 24 hours.