PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026Sensors3 citationsOpen Access

A Lightweight Radar–Camera Fusion Deep Learning Model for Human Activity Recognition

View Full Paper
JYJeanie YooSWSungmin Woo

Key Points

  • The aim is to develop a robust model for recognizing human activities using radar and camera inputs while ensuring user privacy.
  • Developed a radar-camera fusion deep learning model
  • Processed radar data as a Range-Doppler-Time cube
  • Used a privacy-preserving ultra-low-resolution camera
  • Employs Transformer-based models for both radar and camera inputs
  • Constructed a multimodal dataset of synchronized radar and camera sequences
  • Achieved a classification accuracy of 98.74%
  • Outperformed single-modality baselines significantly
  • Model efficiency is demonstrated with only 11 million floating-point operations

Abstract

Human activity recognition in privacy-sensitive indoor environments requires sensing modalities that remain robust under illumination variation and background clutter while preserving user anonymity. To this end, this study proposes a lightweight radar–camera fusion deep learning model that integrates motion signatures from FMCW radar with coarse spatial cues from ultra-low-resolution camera frames. The radar stream is processed as a Range–Doppler–Time cube, where each frame is flattened and sequentially encoded using a Transformer-based temporal model to capture fine-grained micro-Doppler patterns. The visual stream employs a privacy-preserving 4×5-pixel camera input, from which a temporal sequence of difference frames is extracted and modeled with a dedicated camera Transformer encoder. The two modality-specific feature vectors—each representing the temporal dynamics of motion—are concatenated and passed through a lightweight fully connected classifier to predict human activity categories. A multimodal dataset of synchronized radar cubes and ultra-low-resolution camera sequences across 15 activity classes was constructed for evaluation. Experimental results show that the proposed fusion model achieves 98.74% classification accuracy, significantly outperforming single-modality baselines (single-radar and single-camera). Despite its performance, the entire model requires only 11 million floating-point operations (11 MFLOPs), making it highly efficient for deployment on embedded or edge devices.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yoo et al. (2026) studied this question.

synapsesocial.com/papers/6980fd60c1c9540dea80f108https://doi.org/10.3390/s26030894
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Lightweight FMCW radar framework for human activity recognition under limited data conditions2026
  2. 2A Hybrid Millimeter-Wave Radar–Ultrasonic Fusion System for Robust Human Activity Recognition with Attention-Enhanced Deep Learning2026 · 4 citations
  3. 3Radar based continuous indoor activity recognition using deep learning2024 · 1 citations
  4. 4Privacy-Preserving Ambient Sensing for Activities of Daily Living: Multimodal Radar–Thermal Human Activity Recognition and Smart Plug Appliance Recognition2026
  5. 5Synthesizing mmWave range-doppler data from videos for privacy-preserving human activity recognition2026 · 1 citations