Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Appearance Pointers: Multimodal Region Control of Diffusion Transformers

First seen · 7/22/2026, 12:00 PMLatest activity · 7/22/2026, 12:00 PM

This paper introduces Appearance Pointers, compact control tokens for localized multimodal guidance in Diffusion Transformers (DiTs). A region correspondence network aligns text or image inputs with user-provided masks, while spatial aggregation refines the resulting pointers and supports multiple regional descriptions without substantially increasing token load. The method is designed as a modality-agnostic interface that works without retraining the base model from scratch. According to the abstract, a single model matches or surpasses modality-specific state-of-the-art methods across several metrics, although the supplied material does not include the model configuration, benchmark names, numerical results, or implementation details.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/22, 01:59 AMnot independent
    Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
  2. AggregatorHuggingFace Daily Papers7/22, 12:00 PMnot independentRepresentative
    Appearance Pointers: Multimodal Region Control of Diffusion Transformers