Key points are not available for this paper at this time.
This study systematically examines deep learning-based rotoscoping systems along the axes of video object and instance segmentation, memory-based models, optical flow integration, foundation and transformer-based video approaches, as well as matting and production-oriented evaluation metrics. The aim is to classify contemporary rotoscoping methods within a holistic framework, reveal their strengths and limitations under production conditions, and propose a production-oriented evaluation perspective for future research. Methodologically, prominent post 2015 approaches, including VOS/VSS models, memory-based and optical flow-based methods, prompt-driven foundation segmentation models, and transformer-based video systems, are analyzed comparatively with respect to datasets, training strategies, evaluation metrics, and application scenarios. The findings indicate that rotoscoping has evolved from frame-by-frame manual tools toward human-supervised hybrid automation systems based on memory-augmented video segmentation, optical flow-assisted propagation, prompt-based foundation models, and transformer-based video approaches. However, domain gap issues, computational costs in high-resolution sequences, limitations in fine-detail preservation and matting consistency, long-term temporal stability, and the lack of production-specific evaluation metrics remain significant challenges, rendering fully automated rotoscoping an unresolved problem under real-world production conditions. The study suggests that rotoscoping workflows will become highly automated in the near future, yet quality assurance and creative decision-making will continue to rely on human experts within human-in-the-loop hybrid architectures. Accordingly, future research should prioritize standardized evaluation protocols, methods tailored to high-resolution and long-duration video sequences, and hybrid system designs.
Deniz Yuce (Fri,) studied this question.