PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 2026Computer Graphics Forum0 citations

TextFlux: An OCR‐Free DiT Model for High‐Fidelity Multilingual Scene Text Synthesis

View Full Paper
YXYu XieJZJielei ZhangPCPengyu Chen

Key Points

  • To develop a model for multilingual scene text synthesis that eliminates the reliance on OCR and achieves high fidelity.
  • Introduced TextFlux, an OCR-free framework for scene text synthesis.
  • Leveraged diffusion models for contextual reasoning and glyph accuracy.
  • Tested the model’s performance in multilingual settings with low-resource languages.
  • Utilized a streamlined training process requiring only 1% of the data used by other methods.
  • TextFlux outperformed previous methods in both qualitative and quantitative metrics.
  • Achieved effective performance in languages with fewer than 1,000 data samples.
  • Enabled flexible multi-line text generation with precise control.

Abstract

Abstract Diffusion‐based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large‐scale annotated data to support multilingual generation. In this work, we revisit the necessity of complex auxiliary modules and further explore an approach that simultaneously ensures glyph accuracy and achieves high‐fidelity scene integration, by leveraging diffusion models' inherent capabilities for contextual reasoning. To this end, we introduce TextFlux, a DiT‐based framework that enables multilingual scene text synthesis. The advantages of TextFlux can be summarized as follows: (1) OCR‐free model architecture. TextFlux eliminates the need for OCR encoders that are specifically used to extract visual text‐related features. (2) Strong multilingual scalability. TextFlux is effective in low‐resource multilingual settings, and achieves strong performance in newly added languages with fewer than 1,000 samples. (3) Streamlined training setup. TextFlux is trained with only 1% of the training data required by competing methods. (4) Controllable multi‐line text generation. TextFlux offers flexible multi‐line synthesis with precise line‐level control, outperforming methods restricted to single‐line or rigid layouts. Extensive experiments and visualizations demonstrate that TextFlux outperforms previous methods in both qualitative and quantitative evaluations. Our code is available at https://github.com/yyyyyxie/textflux .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xie et al. (2026) studied this question.

synapsesocial.com/papers/69c8c2fcde0f0f753b39d734https://doi.org/10.1111/cgf.70342
Ask AI
Helpful
Bookmark
Share
View Full Paper