Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Are Financial Reasoning Capabilities of LLMs Credible? A Real-World Test over Long-Horizon Statements

First seen · 7/22/2026, 04:00 AMLatest activity · 7/22/2026, 04:00 AM

The paper introduces FinIndices, a benchmark built from uncropped financial statements of up to 32K tokens. It evaluates single-index computation and multi-table index tabulation, including cross-statement, temporal, and stock-flow reasoning. According to the abstract, removing explicit formula hints reduces Gemini-3.1-Pro’s table-task performance from 70.70% to 38.22%, indicating fragile pattern matching and failures in temporal de-cumulation and accounting caliber alignment. Multi-metric, multi-period output also causes models to fall back to shallow heuristics. Supervised fine-tuning improves zero-hint performance by 8.54% on Single-Index and 3.82% on Table-Index tasks.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/22, 04:00 AMnot independentRepresentative
    Are Financial Reasoning Capabilities of LLMs Credible? A Real-World Test over Long-Horizon Statements