Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval

First seen · 7/1/2026, 11:20 AMLatest activity · 7/1/2026, 11:20 AM

This paper introduces FoCo for zero-shot composed image retrieval (ZS-CIR), where a reference image and a textual modification specify the target image. Instead of relying on fixed composition rules such as pseudo-word injection or linear feature arithmetic, FoCo learns composition through two coordinated proxy tasks: text-anchored visual aggregation to focus on modification-relevant content, and context-conditioned semantic completion to form the target representation. Both tasks are jointly trained with a cross-instance contrastive objective intended to encourage semantic diversity and reduce shortcut solutions. The authors report state-of-the-art results and improved generalization across four ZS-CIR benchmarks, although the abstract does not provide benchmark names or numerical gains.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/1, 11:20 AMnot independentRepresentative
    Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval