The paper introduces Goku, a dataset of 2 million high-quality, instruction-aligned video editing pairs covering appearance edits as well as multi-task and structural manipulations such as precise subject-motion control. It also presents Goku-Edit, which uses an MLLM text encoder and a decoupled dual-branch architecture with a dedicated mask branch, plus Goku-Bench, a benchmark containing 1,000 human-verified test cases and seven editing-specific metrics. The authors report up to an 8% instruction-following improvement over other open-source models on Goku-Bench.
No heat snapshots are available in the last 24 hours.