MSEditor: Toward Consistent Multi-Shot Video Editing

Kunyu Feng1, Yue Ma2, Bingyuan Wang1, Yuefeng Wang3, Zhiyuan Qin4, Hao Cheng5, Hao Li4, Qifeng Chen2, Zeyu Wang1,2
The Hong Kong University of Science and Technology (Guangzhou)1, The Hong Kong University of Science and Technology2, Baidu Inc.3, Beijing Innovation Center of Humanoid Robotics4, Tsinghua University5

Abstract

Recent advances in generative AI have significantly improved video generation and editing, achieving impressive results within a single shot. However, extending these capabilities to multi-shot video editing remains a great challenge. Existing methods often suffer from identity drift and cumulative error propagation, where minor inconsistencies in earlier shots amplify over time, leading to unstable subject appearance and degraded visual continuity. Moreover, there is a lack of high-quality multi-shot datasets, further limiting model generalization capabilities. To address these issues, we present \textbf{MSEditor}, the first framework designed for consistent multi-shot video editing. To overcome data scarcity, we repurpose existing multi-view video datasets and introduce a Supervisory Adapter that injects cross-shot information into the backbone, enabling the model to learn identity-consistent representations. Furthermore, we design a Cross-Shot Packing strategy that dynamically aggregates information from related shots, effectively mitigating cumulative errors and enhancing long-range temporal coherence. Extensive experiments demonstrate that MSEditor achieves state-of-the-art performance on multi-shot video editing benchmarks, delivering superior identity preservation, temporal stability, and overall visual quality compared to existing methods.Teaser:

PDF BibTeX
BibTeX copied to clipboard