This paper evaluates whether large language models can distill complete and correct Answer Set Programming theories with a solver in the loop. Starting from one prompt and an empty file, each model has one hour to construct a theory for visual question answering. Across CLEVR, GQA, and CLEVRER, nine models are tested. Three frontier models reach 100% on CLEVR, while frontier-model performance is generally 92.8%–98.8% on GQA. GPT-5 is an outlier, scoring 41.8% on GQA despite 98.7% on CLEVR.
No heat snapshots are available in the last 24 hours.