Excited to share our new work, ๐ฉ๐ถ๐ง๐ฒ๐ซ-๐๐ฒ๐ป๐ฐ๐ต: ๐๐ฒ๐ป๐ฐ๐ต๐บ๐ฎ๐ฟ๐ธ๐ถ๐ป๐ด ๐๐ถ๐ด๐ต-๐๐ถ๐ฑ๐ฒ๐น๐ถ๐๐ ๐ฉ๐ถ๐ฑ๐ฒ๐ผ ๐ฆ๐ฐ๐ฒ๐ป๐ฒ ๐ง๐ฒ๐
๐ ๐๐ฑ๐ถ๐๐ถ๐ป๐ด, accepted to the
#NeurIPS 2026 ๐&๐ ๐ง๐ฟ๐ฎ๐ฐ๐ธ! ๐
Editing text in videos is surprisingly challenging: a good model needs to generate the correct text, keep it temporally consistent across frames, and preserve everything else in the scene.
To study this problem better, we introduce ๐ฉ๐ถ๐ง๐ฒ๐ซ-๐๐ฒ๐ป๐ฐ๐ต, a comprehensive benchmark for video scene text editing. Our release includes:
๐น ๐ฉ๐ถ๐ง๐ฒ๐ซ-๐๐ฎ๐๐ฎ๐๐ฒ๐ โ 387 real-world 720p videos with text-region masks and editing instructions
๐น ๐ ๐๐ต๐ฟ๐ฒ๐ฒ-๐ฎ๐
๐ถ๐ ๐ฒ๐๐ฎ๐น๐๐ฎ๐๐ถ๐ผ๐ป ๐ฝ๐ฟ๐ผ๐๐ผ๐ฐ๐ผ๐น covering text correctness, temporal quality, and edit locality with 13 metrics
๐น ๐๐
๐๐ฒ๐ป๐๐ถ๐๐ฒ ๐ฒ๐๐ฎ๐น๐๐ฎ๐๐ถ๐ผ๐ป ๐ผ๐ณ ๐ด ๐ฏ๐ฎ๐๐ฒ๐น๐ถ๐ป๐ฒ๐ across four video-editing paradigms
๐น ๐ฉ๐ถ๐ง๐ฒ๐ซ-๐๐ฑ๐ถ๐-๐ญ๐ฐ๐, an open-source reference model with motion-aligned glyph-video conditioning
๐น Open-source ๐ฑ๐ฎ๐๐ฎ, ๐ฒ๐๐ฎ๐น๐๐ฎ๐๐ถ๐ผ๐ป ๐ฐ๐ผ๐ฑ๐ฒ, ๐บ๐ผ๐ฑ๐ฒ๐น ๐๐ฒ๐ถ๐ด๐ต๐๐, ๐ฎ๐ป๐ฑ ๐น๐ฒ๐ฎ๐ฑ๐ฒ๐ฟ๐ฏ๐ผ๐ฎ๐ฟ๐ฑ
One takeaway from our study: no single metric tells the whole story. Models that preserve the scene well may fail to actually edit the text, while models that generate the right text frame-by-frame can still suffer from severe temporal flickering.
ViTeX-Bench is designed to make these trade-offs explicit and reproducible. We hope ViTeX-Bench can provide a useful foundation for future research on precise and temporally consistent video editing.
๐ Project page:
vitex-bench.github.io
๐ Paper link:
arxiv.org/abs/2609.40356
#NeurIPS2026 #ComputerVision #GenerativeAI #VideoEditing #VideoGeneration #MultimodalAI #MachineLearning #AIResearch