The paper introduces Graph-PRefLexOR, a family of graph-native reasoning models trained with Group Relative Policy Optimization (GRPO). Its reasoning process is organized into mechanism exploration, graph construction, pattern extraction, and hypothesis synthesis. On 100 open-ended questions drawn from materials science and mechanics literature, the authors report 40–65% improvements over corresponding base models, with the largest gains in traceability. Embedding analyses indicate roughly 2–3 times greater semantic diversity. Test-time graph expansion mainly improves long-range conceptual recombination within a bounded semantic space rather than continually broadening semantic coverage.
No heat snapshots are available in the last 24 hours.