MSEditor: Toward Consistent Multi-Shot Video Editing
Abstract
Recent advances in generative AI have significantly improved video generation and editing, achieving impressive results within a single shot. However, extending these capabilities to multi-shot video editing remains a great challenge. Existing methods often suffer from identity drift and cumulative error propagation, where minor inconsistencies in earlier shots amplify over time, leading to unstable subject appearance and degraded visual continuity. Moreover, there is a lack of high-quality multi-shot datasets, further limiting model generalization capabilities. To address these issues, we present \textbf{MSEditor}, the first framework designed for consistent multi-shot video editing. To overcome data scarcity, we repurpose existing multi-view video datasets and introduce a Supervisory Adapter that injects cross-shot information into the backbone, enabling the model to learn identity-consistent representations. Furthermore, we design a Cross-Shot Packing strategy that dynamically aggregates information from related shots, effectively mitigating cumulative errors and enhancing long-range temporal coherence. Extensive experiments demonstrate that MSEditor achieves state-of-the-art performance on multi-shot video editing benchmarks, delivering superior identity preservation, temporal stability, and overall visual quality compared to existing methods.Teaser:
