Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

CAPEval: Decoupled Caption Evaluation for Understanding and Generation

First seen · 8/3/2026, 04:00 AMLatest activity · 8/3/2026, 04:00 AM

CAPEval proposes evaluating image captions through two separate dimensions: Coverage, which measures how thoroughly a caption captures factual visual content, and Precision, which measures the fraction of its claims that are supported by the image. The benchmark uses human-written ground-truth captions and human-verified atomic checklist items. Experiments compare 10 captioners from four model families under controlled downstream settings where caption source is the only changing variable. The reported results show that Coverage is more strongly associated with multimodal understanding, while Precision is the dominant predictor of text-to-image generation performance.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/3, 04:00 AMnot independentRepresentative
    CAPEval: Decoupled Caption Evaluation for Understanding and Generation