The paper presents Pythia, a multi-agent system that autonomously writes and optimizes prompts for extracting clinical signs and symptoms without manual prompt engineering or fine-tuning. Running on a locally hosted open-weights model, it was evaluated on 400 clinical notes covering 72 concepts. Pythia achieved mean sensitivity of 0.76 and specificity of 0.95, outperforming a curated lexicon on specificity and recovering specificity for concepts whose lexicon marked every note positive. However, sensitivity transfer weakened substantially for rare concepts, especially below 2% prevalence.
No heat snapshots are available in the last 24 hours.