The paper introduces GRAPHEVAL, a graph-based framework that evaluates LLM reasoning as a structured space rather than relying only on final-answer agreement. Its Graph Reasoning Coherence Score (GRCS) measures semantic and structural consensus, aiming to expose mode collapse and confident hallucination. Graph Self-Consistency (GSC) selects a medoid reasoning path instead of using naive majority voting. The abstract reports that GRCS is consistently negatively correlated with reasoning faithfulness across larger and smaller models, while adversarial medoid ablation suggests that the selected path can be load-bearing for both faithfulness and, in targeted cases, accuracy.
No heat snapshots are available in the last 24 hours.