LeapTalk introduces a single-step bridge-distillation framework for long-form and real-time talking-head generation. It replaces the conventional noise-to-data formulation with data-to-data transport based on a Brownian bridge, anchored by a persistent reference to reduce identity drift and improve long-term stability. The method also uses an SNR-aligned time transformation to transfer knowledge from a pretrained diffusion teacher to a heterogeneous student model, plus audio-driven classifier-free guidance for lip synchronization. The authors report one-step generation at up to 200 FPS and claim improved efficiency and temporal consistency, although detailed benchmarks are not included in the abstract.
No heat snapshots are available in the last 24 hours.