The paper introduces Perceptual Flow Matching (PFM), which performs flow-matching supervision in the feature space of pretrained perceptual models rather than the conventional VAE latent space. According to the abstract, this change reduces sampling from 35–50 steps to 4–8 while preserving generation quality across image generation, video generation, and image editing. PFM does not require a teacher model or auxiliary score network and can be added to standard flow-matching pipelines with minimal changes. The authors attribute the gain to a shift from mean-seeking to mode-seeking regression, producing predictions that are more robust to coarse few-step integration.
No heat snapshots are available in the last 24 hours.