This approach demonstrates improved font style transfer in scene text editing, suggesting robust evaluation metrics.
We present a novel approach for font style transfer using Generative Adversarial Networks (GANs) to enhance the scene text editing process, enabling text editing from any characters to any characters, including cross-language editing. Our GAN model utilizes pairs of sample images of the target font style and the corresponding skeleton-based features to learn their key structural details without relying on pre-trained models. Once the generator is trained, it can transform any character from the base font style to the target font style. Our approach offers the flexibility to select a base font similar to the target font for enhancing results and the ability to manipulate the stroke width of the output text. Additionally, in few-shot scenarios, we introduce a double generator scheme that integrates other existing methods with our approach. In this work, we also introduce two new evaluation metrics: Difference in Histogram of Oriented Gradients and Stroke Width Similarity. Our experimental results demonstrate that the proposed evaluation metrics can better measure font style similarity with greater robustness compared to conventional metrics. We evaluate the performance of our GAN model on style transfer for six target fonts and real scene text editing tasks, comparing it with existing methods. Our approach provides better structural similarity, readability, and visual appeal than other methods, especially for generating unseen characters.
No takes yet. Share an insight, caveat, or question.
Thanusan et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: