OpenAI describes an API configuration experiment for GPT-5.6 on the ARC-AGI-3 benchmark. According to the supplied summary, retaining reasoning and enabling compaction tripled the model’s score while improving efficiency. The available material does not provide baseline or final absolute scores, token or cost measurements, task-level results, or independent replication. The result therefore points to a potentially important inference-configuration effect, but its magnitude and generality cannot yet be assessed from the supplied information alone.
No heat snapshots are available in the last 24 hours.