PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 13, 2026Humanities and Social Sciences Communications0 citationsOpen Access

Using vision-language models to extract network data from images of system maps

View Full Paper
JWJordan WhiteUniversity of OxfordPBPete Barbrook-JohnsonUniversity of Oxford

Key Points

  • The research aims to evaluate the effectiveness of vision-language models in extracting data from system maps.
  • Examined seven vision-language models for information extraction from images of system maps.
  • Tested on three types of system map diagrams: Causal Loop Diagrams, Fuzzy Cognitive Maps, and Theory of Change diagrams.
  • Assessed extraction quality using formats: DOT, JSON, and Markdown table.
  • Models summarize factors in maps better than connections.
  • Some models perfectly extract factor labels for specific images and formats.
  • Models perform better with visually distinct diagrams and consistent node-edge relationships.

Abstract

Abstract A range of systems mapping approaches are widely used to support the analysis and design of public policy, but can be time and resource intensive to implement. Generative AI tools may be able to streamline the use of systems mapping by helping researchers to quickly synthesise existing data on policy systems, freeing resources to foster greater stakeholder participation and use of maps. To explore and test the potential of these tools to help with systems mapping exercises, we examine the performance of seven proprietary vision language models (VLMs) with a key task in potential workflows - extraction of relevant information from images of system maps already created. VLMs present value as they allow for the synthesis of both textual and image data simultaneously. We test on images of three types of system map diagrams: Causal Loop Diagrams, Fuzzy Cognitive Maps, and Theory of Change diagrams, and test three different formats for structuring data: DOT, JSON and Markdown table. We find that models summarise factors in maps better than connections, with some models extracting factor labels perfectly for certain images and formats. Models appear to perform better with diagrams that have bolder graphics and when there is greater internal consistency between separate node and edge lists. We also find that models appear to omit correct information more than they include false information, although falsehoods are still common. Our formal approach to testing introduces an empirical framework that will allow researchers to conduct similar research in the future, to maintain pace as the application and capabilities of language models continue to evolve.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

White et al. (2026) studied this question.

synapsesocial.com/papers/6a03cbe01c527af8f1ecfa04https://doi.org/10.1057/s41599-026-07537-w
Ask AI
Helpful
Bookmark
Share
View Full Paper