PoTRE, or Poly-Topological Reasoning Ensembles, is a test-time reasoning framework that separates inference into four heterogeneous agents: adversarial refinement, hierarchical strategic planning, spectrum search, and direct chain reasoning. A task-adaptive aggregation layer reconciles their outputs through candidate selection, semantic synthesis, or neuro-symbolic verification. The paper evaluates the method on ARC-AGI-2, Humanity's Last Exam (HLE), and PRBench Finance. It reports 49.92% accuracy on HLE, described as exceeding the previous best official score, while using similar or fewer inference tokens than heavily scaled homogeneous baselines.
No heat snapshots are available in the last 24 hours.