This project proposes a methodology for aligning vector models to improve labor market analysis across Europe, highlighting the importance of model selection.
Over the last decade, word embeddings have enabled machines to represent words and sentences as vectors, enabling researchers to reason on text for tasks like semantic similarity, contextual understanding, machine translation, etc. However, the synthesis of embeddings involves domain-specific parameters that affect semantic accuracy and contextual relevance, often leading to unpredictable biases and inconsistent comparisons. This issue is particularly relevant in labor market analysis, where different embeddings yield varying results, making the selection of the most appropriate model a key element. This paper addresses these challenges by (i) proposing a methodology to train, select, and align vector space models for a target taxonomy, ensuring comparability across dimensions and languages; (ii) applying this approach to 4.5 million job ads in 28 languages, aligning country-specific embeddings using the ESCO taxonomy; (iii) generating over 3000 models over 142 machine days, making the best-performing ones publicly available via VEUCTOR ; and (iv) showing how model choice significantly impacts labor market analysis, revealing substantial variations in occupational skill bundles across embeddings. • We present, formalise, and implement a multilingual methodology to train, select, and align word embedding models using the ESCO taxonomy across 28 European countries. • We generate and evaluate over 3000 embedding models trained on 4.5 million online job advertisements in the frame of an EU Project, using a benchmark-driven approach to optimize semantic alignment. • We release VEUCTOR , a tool that provides access to the best-performing and aligned embeddings, enabling reuse and supporting third-party labor market analyses. • We show that the choice of embedding significantly affects occupational skill bundles and, consequently, labor market analysis outcomes. • We enable reproducible and cross-country labor market intelligence by standardizing model development and alignment across diverse languages and corpora.
No takes yet. Share an insight, caveat, or question.
Colombo et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: