LongE2V fine-tunes a pretrained video diffusion model for three sparse-event-stream tasks: video reconstruction, future prediction, and frame interpolation. The method combines Autoregressive Unrolling and Adaptive Context Switching to reduce temporal drift over very long sequences, while Reencoding Alignment with Cross Residual Correction targets bidirectional consistency during interpolation. Event Voxel Density Augmentation is intended to improve robustness across sensor resolutions. The abstract reports state-of-the-art results on real-world benchmarks and zero-shot generalization, but detailed metrics, baselines, and evaluation protocols require inspection of the full paper.
No heat snapshots are available in the last 24 hours.