FlowMimic presents a pixel-pair temporal warped flow field that generates corresponding video-editing samples from image-editing samples in real time, reducing reliance on object masks, I2V-based pair synthesis, ControlNet-like guidance, and VLM filtering. The model treats images as a special case of video and uses modality-mimic generation and editing losses to align the two modalities. It also adds sense-related tasks, including referring expression segmentation, together with editing-region-aware latent and attention losses, aiming to internalize instruction understanding, region localization, and localized modification without mask sequences or an additional MLLM.
No heat snapshots are available in the last 24 hours.