PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 8, 2026Machine Learning and Knowledge Extraction0 citationsOpen Access

Enhancing the Extraction of GHG Emission-Reduction Targets from Sustainability Reports Using Vision Language Models

LWLars WilhelmiCBChristian BrunsMSMatthias Schümann

Key Points

  • The central aim is to enhance the extraction of GHG emission-reduction targets from sustainability reports using VLMs.
  • Designed an extraction artifact using Design Science Research Methodology.
  • Developed a curated page-level dataset of GHG emission-reduction targets.
  • Implemented an automated evaluation pipeline comparing various model inputs and preprocessing.
  • Conducted a controlled comparison of text, image, and combined modalities in extraction.
  • Assessed extraction quality using F1-scores at the page and attribute level.
  • Mistral Small 3.2 showed the most stable performance in extracting metrics.
  • Combined text and image modality achieved the highest F1 score of 0.82, especially with complex layouts.
  • Visual–textual structure preserved led to improved extraction quality compared to text-only methods.
  • Challenges were noted for visually dense layouts and preventing inference-based hallucinations.

Abstract

This study investigates how Vision Language Models (VLMs) can be used and methodically configured to extract Environmental, Social, and Governance (ESG) metrics from corporate sustainability reports, addressing the limitations of existing text-only and manual ESG data-extraction approaches. Using the Design Science Research Methodology, we developed an extraction artifact comprising a curated page-level dataset containing greenhouse gas (GHG) emission-reduction targets, an automated evaluation pipeline, model and text-preprocessing comparisons, and iterative prompt and few-shot refinement. Pages from oil and gas sustainability reports were processed directly by VLMs to preserve visual–textual structure, enabling a controlled comparison of text, image, and combined input modalities, with extraction quality assessed at page and attribute level using F1-scores. Among tested models, Mistral Small 3.2 demonstrated the most stable performance and was used to evaluate image, text, and combined modalities. Combined text + image modality performed best (F1 = 0.82), particularly on complex page layouts. The findings demonstrate how to effectively integrate visual and textual cues for ESG metric extraction with VLMs, though challenges remain for visually dense layouts and avoiding inference-based hallucinations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wilhelmi et al. (2026) studied this question.

synapsesocial.com/papers/6988291e0fc35cd7a884932ahttps://doi.org/10.3390/make8020037
Ask AI
Helpful
Bookmark
Share
View Full Paper