This paper presents a training-free method for fine-grained identity tuning in text-to-image personalization models. Instead of editing a single input image, it modifies the latent representation of a target identity, allowing the edited identity to be generated across diverse images. The method analyzes a pretrained, frozen personalization encoder and identifies semantic directions in its latent token space, including token-defined subspaces associated with particular facial or spatial regions. These directions support localized and semantically coherent edits while preserving identity consistency across generated images. The authors report qualitative and quantitative validation, but the provided abstract does not include benchmark names or numerical results.
No heat snapshots are available in the last 24 hours.