PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 12, 2026Discover Computing0 citationsOpen Access

A Hybrid deep learning framework for capturing environmental change through image captioning

MMMaryam MehmoodAIAsad IjazNANouman Ali

Key Points

  • The aim is to create an automatic image captioning system for remote sensing images to support environmental monitoring.
  • Developed a hybrid framework combining visual features from VGG-16 and semantic representations from Word2Vec.
  • Utilized an attention-enhanced Long Short-Term Memory (LSTM) network for decoding and generating captions.
  • Applied the framework to UC Merced Land Use (UCM) and RSICD datasets to evaluate environmental categories.
  • Produced high-accuracy captions describing land use and ecological conditions.
  • Demonstrated superior performance against existing methods using evaluation metrics like BLEU and METEOR.
  • Captured critical environmental attributes relevant for climate change and sustainable planning.

Abstract

Remote sensing plays a central role in monitoring Earth’s surface for environmental changes such as deforestation, urban expansion, water scarcity, and climate-induced disasters. However, the rapid increase in satellite image acquisition makes manual interpretation impractical. This study proposes a hybrid deep learning framework that automatically generates descriptive captions for remote sensing images, enabling environmental scientists to interpret large-scale Earth observation data efficiently. The framework integrates visual features extracted with a fine-tuned VGG-16 network and semantic representations learned through Word2Vec embeddings, which are fused and decoded via an attention-enhanced Long Short-Term Memory (LSTM) network. Applied to the UC Merced Land Use (UCM) and RSICD datasets, which cover diverse environmental categories, the model produces captions that describe land use and ecological conditions with high accuracy. Evaluation using BLEU, METEOR, ROUGE, and CIDEr metrics demonstrates superior performance compared to existing approaches. More importantly, the generated captions capture meaningful environmental attributes–such as vegetation loss, settlement growth, or presence of water bodies–that are critical for applications in climate change monitoring, disaster management, and sustainable land-use planning. This approach provides a pathway for large-scale, automated environmental assessments, supporting decision-making in Earth system science and policy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mehmood et al. (2026) studied this question.

synapsesocial.com/papers/69b2582a96eeacc4fcec7788https://doi.org/10.1007/s10791-026-10025-z
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Remote Sensing Image Captioning Using Deep Learning2024 · 3 citations
  2. 2Remote Sensing Image Captioning via Self-Supervised DINOv3 and Transformer Fusion2026 · 1 citations
  3. 3Understanding remote sensing imagery like reading a text document: What can remote sensing image captioning offer?2024 · 4 citations
  4. 4Remote Sensing Image Captioning Using Transformer Model and CNN Feature Extraction Model2024
  5. 5SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning2025