PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 15, 2026Information0 citationsOpen Access

HFI-Former: High-Frequency Interaction Transformer for Robust Scene Text Detection

View Full Paper
YGYubing GaoQGQuanli GaoLSLianhe Shao

Key Points

  • The aim is to improve scene text detection, especially in complex environments with cluttered backgrounds.
  • Developed a Transformer-based model called HFI-Former.
  • Implemented multi-scale feature extraction to capture various detail levels.
  • Introduced frequency-domain enhancement to protect high-frequency features from degradation.
  • Utilized semantic-aware feature interaction for better context regulation in feature fusion.
  • Achieved competitive boundary localization accuracy on three datasets: CTW1500, Total-Text, and ICDAR1500.
  • Showed strong overall text detection performance in complex scenes.

Abstract

Scene text detection aims to accurately localize text instances in images captured under complex environments. Its performance depends heavily on precise text boundary delineation and reliable semantic discrimination from cluttered backgrounds. However, existing methods still struggle in such complex scenes. Repeated downsampling gradually biases features toward low-frequency components, thereby weakening edge details and local structures that are critical to text morphology. Additionally, semantic information and local details are often modeled independently. This lack of coordination makes high-frequency responses vulnerable to background noise. To address these issues, we propose HFI-Former, a Transformer-based model designed for high-frequency enhancement and feature interaction. The framework consists of multi-scale feature extraction, frequency-enhanced representation, semantic-guided feature interaction, and deformable Transformer encoding. Frequency-domain enhancement is introduced to preserve high-frequency structural features degraded by repeated downsampling. Semantic-aware feature interaction further injects global context to regulate multi-scale feature fusion. Experiments on CTW1500, Total-Text and ICDAR1500 demonstrate competitive boundary localization accuracy and strong overall detection performance in complex scenes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gao et al. (2026) studied this question.

synapsesocial.com/papers/69df2c2fe4eeef8a2a6b134dhttps://doi.org/10.3390/info17040365
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A Multi-Scale Natural Scene Text Detection Method Based on Attention Feature Extraction and Cascade Feature Fusion2024 · 10 citations
  2. 2LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation Network2024 · 20 citations
  3. 3Focus Entirety and Perceive Environment for Arbitrary-Shaped Text Detection2024 · 13 citations
  4. 4ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer2023 · 42 citations
  5. 5CBNet: A Plug-and-Play Network for Segmentation-Based Scene Text Detection2024 · 17 citations