The paper introduces FinIndices, a benchmark built from uncropped financial statements of up to 32K tokens. It evaluates single-index computation and multi-table index tabulation, including cross-statement, temporal, and stock-flow reasoning. According to the abstract, removing explicit formula hints reduces Gemini-3.1-Pro’s table-task performance from 70.70% to 38.22%, indicating fragile pattern matching and failures in temporal de-cumulation and accounting caliber alignment. Multi-metric, multi-period output also causes models to fall back to shallow heuristics. Supervised fine-tuning improves zero-hint performance by 8.54% on Single-Index and 3.82% on Table-Index tasks.
No heat snapshots are available in the last 24 hours.