This paper identifies “exception chain collapse,” a failure mode in which frontier LLMs mis-evaluate nested eligibility rules such as “A is required unless B applies, unless C overrides B.” It reports silent changes in failure cells under the same model alias and prompt, including GPT-5.4 improving from 96.6% to 100% on a construction-insurance task between March and April 2026. The proposed Aethis Eligibility Module uses LLMs to author rules from authoritative sources, then executes them deterministically with an SMT-based layer. The abstract reports results on 225 benchmark scenarios, 20 adversarial cases, and 949 LegalBench cases.
No heat snapshots are available in the last 24 hours.