The paper proposes Dualin, a two-stage inversion method that jointly recovers a human-interpretable hard prompt and the latent noise associated with a target image. The first stage combines a vision-language model, CLIP, and a large language model to produce a faithful prompt. The second applies unconditional DDIM inversion to reconstruct the target’s latent noise, preserving structural information. The authors claim state-of-the-art image fidelity across diverse datasets and argue that the recovered noise supports flexible editing without re-optimization.
No heat snapshots are available in the last 24 hours.