Observational analysis highlights performance comparison of LLMs on systems engineering tasks, suggesting necessary benchmarks.
In the systems engineering (SE) community, generative artificial intelligence (GenAI), such as large language models (LLMs), continues to grow in popularity with applications varying from help with writing requirements, to creating traditional text‐based documents, and to generating models. Furthermore, the text‐based format of the Systems Modeling Language version 2 (SysMLv2) gives rise to wide ranging possibilities when combined with LLMs. It is common to specialize LLMs on specific tasks, such as SE; which may be achieved by a number of means. Fine‐tuning and retrieval augmented generation (RAG) are example methods to teach a model how to act and what the model needs to know, respectively. Typical practice evaluates LLMs relative to one‐another through the use of benchmarks. Benchmarks are tasks that use a common scoring‐based method where the LLM either answers a set of multiple‐choice questions or textual queries to verified, correct answers. A domain specific benchmark for SE, SysEngBench, was developed over the last year. With many in the SE community developing custom language models, the usage of a standard benchmark is essential for relative performance comparison. We have conducted a limited experiment using SysEngBench to compare performance, on SE tasks, of a small set of LLMs (control plus two modified models). The intent of this study was to set a baseline for ranking models according to performance on SE tasks, the results suggest what many in the SE community may find intuitively unexpected.
No takes yet. Share an insight, caveat, or question.
Wach et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: