This paper introduces Appearance Pointers, compact control tokens for localized multimodal guidance in Diffusion Transformers (DiTs). A region correspondence network aligns text or image inputs with user-provided masks, while spatial aggregation refines the resulting pointers and supports multiple regional descriptions without substantially increasing token load. The method is designed as a modality-agnostic interface that works without retraining the base model from scratch. According to the abstract, a single model matches or surpasses modality-specific state-of-the-art methods across several metrics, although the supplied material does not include the model configuration, benchmark names, numerical results, or implementation details.
No heat snapshots are available in the last 24 hours.