The paper introduces CRTBench, a benchmark with 350 question families and 1,750 questions, designed to test whether language models preserve answers under controlled logically equivalent reformulations. Transformations include contrapositives, double negation, negation flipping, and passive voice. GPT-5.4-mini reaches 98.9% base accuracy but only 60.3% family-level consistency, while reasoning-optimized o4-mini reaches 96.9% consistency. Performance is weakest on contrapositives and double negation, whereas surface paraphrases remain comparatively robust. Increasing reasoning effort raises GPT-5.4-mini consistency to 85.4%, but does not improve GPT-5.4 overall because gains on nested negation are offset by quantifier-family failures.
No heat snapshots are available in the last 24 hours.