Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 15, 2026Proceedings of the Estonian Academy of SciencesOpen Access

The efficiency–verification trade-off in large language model-assisted workplace risk assessment

View Full Paper
Ask AI
Bookmark
Share

Authors

ABAleksandr BoslerTKTarmo KoppelTallinn University of TechnologyKRKarin ReinholdTallinn University of Technology

Discussion

Loading...

Member takes

Implication

Evaluation study demonstrates significant time savings but notable AI-human disagreement in student risk assessments, highlighting potential overreliance on language models.

Key Points

  • To evaluate the trade-off between perceived efficiency gains and the verification burden when applying large language models to workplace risk assessment workflows.
  • Analyzed 96–121 analytic sessions derived from 111 student assignment files in an educational setting.
  • Operationalized efficiency via perceived completion time versus an estimated manual baseline, and verification burden as AI–human quantitative score disagreement (1–5 scale for probability and impact).
  • Evaluated user acceptance with utility and satisfaction Likert scales and modeled satisfaction predictors using robust ordinary least squares (OLS) regression.
  • LLM assistance reduced median task time from 120 minutes to 20 minutes, representing a median perceived time saving of 0.84.
  • The median AI–human disagreement proxy was 0.50, with 26.8% of sessions having disagreement ≥1.0, and 40.6% of sessions displaying high satisfaction alongside high disagreement.
  • Robust OLS regression showed utility as the strongest positive predictor of satisfaction and disagreement as a negative predictor, whereas time savings were not statistically significant.

Cite This Study

Bosler et al. (2026) studied this question.

synapsesocial.com/papers/6a8019bb75c2e31742c85d88https://doi.org/10.3176/proc.2026.3.10
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1How good are large language models at product risk assessment?2024 · 27 citations
  2. 2The Impact of Performance Expectancy, Workload, Risk, and Satisfaction on Trust in ChatGPT: Cross-Sectional Survey Analysis2024 · 66 citations
  3. 3Automation Bias in AI-Decision Support: Results from an Empirical Study2024 · 50 citations
  4. 4Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI2024 · 413 citations
  5. 5Can ChatGPT exceed humans in construction project risk management?2024 · 73 citations