Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What they did and found

Researchers tested eight ways to edit text in 157 real videos with readable text. In 120-frame clips, drift, flicker, or unchanged text occurred in many cases. They used 230 training videos and 157 test videos, all with Latin letters, to compare text accuracy, image quality, and how well edits stayed on the moving surface.

Why it matters

To avoid growing text errors as scenes change, check the whole clip. A sample shop video should look correct from start to finish.

The ViTeX dataset holds many source clips and 230 training pairs. Edits often lose legibility as motion, lighting, or occlusions change, making glyphs drift.

A mixed method that lines up the letters to the moving surface can help, but it may not stop drift entirely.

What remains uncertain

Limited to Latin-script text and real-world footage; results may vary with other scripts, fonts, or heavy motion.

Read the paper PDF ↗

Original sources · 1
  1. ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing ↗arXiv · 2026-09-30

Check the original paper for its authors, methods, version and access terms.