PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 21, 2026ACM Transactions on Multimedia Computing Communications and Applications1 citations

RCAENet: Residual Convolutional and Attention-Enhanced Stereo Matching for Real-Time Depth Estimation on Edge Devices

View Full Paper
BLBifa LiangYWYichao WangZHZiyang Hu

Key Points

  • The aim is to develop a high-performance stereo network optimized for real-time depth estimation on edge devices.
  • Introduced Residual Convolutional Feature Extraction (RCFE) for efficient feature capture.
  • Developed Enhanced Adaptive Upsampling (EAU) module for improved feature fusion.
  • Designed Enhanced 3D CNN (E3DC) and Cost Aggregation and Residual Attention (CA-ResAgg) module for cost volume regularization.
  • Implemented a multi-scale architecture to balance accuracy and efficiency.
  • Achieved real-time inference on power-constrained edge devices.
  • Maintained state-of-the-art depth accuracy.
  • Enhanced disparity estimation accuracy through innovative feature extraction and attention mechanisms.

Abstract

As a core technology in real-time video processing and intelligent surveillance, stereo matching provides essential depth perception capabilities for multimedia applications. However, high-precision stereo networks often come with significant computational costs, making real-time inference on power- and memory-constrained edge devices challenging. On the other hand, lightweight real-time networks still struggle with accuracy limitations. To address this challenge, we propose RCAENet, a high-performance stereo network designed for real-time and high-accuracy depth estimation on edge devices. To enhance feature extraction efficiency, we introduce the Residual Convolutional Feature Extraction (RCFE) module, which replaces conventional convolutional layers to capture more expressive features while maintaining computational efficiency. Additionally, we propose the Enhanced Adaptive Upsampling (EAU) module, which integrates channel and spatial attention mechanisms to improve feature fusion and disparity refinement. Furthermore, we design an Enhanced 3D CNN (E3DC) along with the Cost Aggregation and Residual Attention (CA-ResAgg) module for cost volume regularization. This module incorporates residual aggregation and efficient channel attention to further enhance disparity estimation accuracy. Built upon these components, RCAENet features a multi-scale architecture that effectively balances accuracy and efficiency. Extensive experiments demonstrate that these innovations enable RCAENet to achieve real‑time inference on edge devices while maintaining state‑of‑the‑art depth accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liang et al. (2026) studied this question.

synapsesocial.com/papers/69706c09b6488063ad5c17c2https://doi.org/10.1145/3788678
Ask AI
Helpful
Bookmark
Share
View Full Paper