Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Controlled Reformulation Testing for Logical Consistency in Large Language Models

First seen · 7/16/2026, 11:30 AMLatest activity · 7/16/2026, 11:30 AM

The paper introduces CRTBench, a benchmark with 350 question families and 1,750 questions, designed to test whether language models preserve answers under controlled logically equivalent reformulations. Transformations include contrapositives, double negation, negation flipping, and passive voice. GPT-5.4-mini reaches 98.9% base accuracy but only 60.3% family-level consistency, while reasoning-optimized o4-mini reaches 96.9% consistency. Performance is weakest on contrapositives and double negation, whereas surface paraphrases remain comparatively robust. Increasing reasoning effort raises GPT-5.4-mini consistency to 85.4%, but does not improve GPT-5.4 overall because gains on nested negation are offset by quantifier-family failures.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/16, 11:30 AMnot independentRepresentative
    Controlled Reformulation Testing for Logical Consistency in Large Language Models