PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 20260 citationsOpen Access

Identifying Sustainability in Public Tendering

View Full Paper
LRLuca Sven RolshovenVMVeton MatoshiTETilia Ellendorff

Key Points

  • The aim is to enhance assessment of sustainability integration in public procurement documents.
  • Developed a Natural Language Processing pipeline to identify sustainability criteria.
  • Compiled four catalogs of Sustainable Procurement Criteria for analysis.
  • Calculated cosine similarity scores for sentence-criterion pairs.
  • Validated similarity threshold using expert reviews and Large Language Models.
  • A similarity threshold of 0.98 effectively identified relevant sustainability criteria.
  • Gemini 2.0 showed substantial agreement with expert judgments (Fleiss’ Kappa = 0.754).
  • LLM-based validation offers a scalable alternative but with variable performance.

Abstract

Public procurement serves as a significant lever for promoting sustainability, yet effectively assessing the integration of sustainability criteria within diverse and heterogeneous tender documents remains a challenge. This paper presents a Natural Language Processing (NLP) pipeline for automatically identifying sustainability criteria in Swiss public procurement documents written in German. To assess sustainability, we compiled four catalogs of official Sustainable Procurement Criteria (SPCs): three domain-specific (transport, food, furniture) and one domain-independent. Each call for tenders (CFT) document was segmented into sentences and encoded using a pre-trained sentence transformer. We then computed cosine similarity scores between each sentence and all SPCs, storing the top match from both the general and the domain-specific catalog, if applicable. While similarity scores were generally high for a majority of sentences, a preliminary manual inspection suggested that only matches with a score of 0.98 or higher tended to reflect meaningful alignment. To validate this threshold, two human experts independently reviewed 100 randomly sampled sentence-criterion pairs above this threshold. To explore whether this expert validation process could be scaled, we also prompted three different Large Language Models (LLMs) to assess the same samples, classifying each pair as a correct or incorrect match based on a majority vote. Our evaluation suggests that a similarity threshold of 0.98 is useful for reducing noise and identifying relevant sustainability criteria. LLM-based validation shows potential as a scalable alternative to human annotation, although performance varies between models. While Gemini 2.0 achieved substantial agreement with the expert judgments in terms of Fleiss’ Kappa (𝜅 = 0.754), other models demonstrated weaker alignment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rolshoven et al. (2026) studied this question.

synapsesocial.com/papers/69b4ad8d18185d8a39801099https://doi.org/10.24451/arbor.13428
Ask AI
Helpful
Bookmark
Share
View Full Paper