ESTONIAN ACADEMY
PUBLISHERS
eesti teaduste
akadeemia kirjastus
PUBLISHED
SINCE 1952
 
Proceeding cover
proceedings
of the estonian academy of sciences
ISSN 1736-7530 (Electronic)
ISSN 1736-6046 (Print)
Impact Factor (2024): 0.7

Research article
The efficiency–verification trade-off in large language model-assisted workplace risk assessment; pp. 251–259
PDF | https://doi.org/10.3176/proc.2026.3.10

Authors
Aleksandr Bosler ORCID Icon, Tarmo Koppel, Karin Reinhold ORCID Icon
Abstract

Large language models (LLMs) are increasingly adopted in occupational health and safety (OHS) to accelerate workplace risk assessment, yet their integration raises a dilemma between efficiency and verification. The study examines this trade-off in an LLM-assisted workflow using data from 111 student assignment files from an educational setting, analyzed at the session level, yielding 96‒121 analytic sessions depending on measure availability. Efficiency was measured via perceived time savings versus an estimated manual baseline, while the verification-burden proxy was operationalized as an AI‒human disagreement in quantitative scoring (probability and impact on 1‒5 scales). User acceptance was captured via the utility and satisfaction / intention to reuse Likert scales.

The results show substantial perceived efficiency gains: the median LLM-assisted time was 20 minutes versus the median manual time of 120 minutes (median saving 0.84). The verification-burden proxy remained non-trivial: median disagreement was 0.50, and 26.8% of sessions exhibited disagreement ≥1.0. Acceptance remained high (mean satisfaction 4.06/5). Notably, 40.6% of the sessions showed a possible overreliance pattern, with high satisfaction coexisting with high disagreement. In a robust ordinary least squares (OLS) regression, utility was the strongest predictor of satisfaction, while disagreement reduced satisfaction; time saving was not significant.

Perceived speed gains may thus coexist with hidden assurance costs and possible trust miscalibration, highlighting the need for verification structures and targeted training for sustainable digital safety governance.

References

Aladağ, H. 2023. Assessing the accuracy of ChatGPT use for risk management in construction projects. Sustainability15(22), 16071.
https://doi.org/10.3390/su152216071

Choudhury, A. and Shamszare, H. 2024. The impact of performance expectancy, workload, risk, and satisfaction on trust in ChatGPT: cross-sectional survey analysis. JMIR Human Factors11(1), e55399.
https://doi.org/10.2196/55399

Collier, Z. A., Gruss, R. J. and Abrahams, A. S. 2024. How good are large language models at product risk assessment? Risk Analysis45(4), 766–789.
https://doi.org/10.1111/risa.14351

Falegnami, A., Tomassi, A., Corbelli, G., Nucci, F. S. and Romano, E. 2024. A generative artificial-intelligence-based workbench to test new methodologies in organisational health and safety. Applied Sciences14(24), 11586.
https://doi.org/10.3390/app142411586

Goddard, K., Roudsari, A. and Wyatt, J. C. 2012. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association19(1), 121–127.
https://doi.org/10.1136/amiajnl-2011-000089

Klingbeil, A., Grützner, C. and Schreck, P. 2024. Trust and reliance on AI ‒ an experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior160, 108352.
https://doi.org/10.1016/j.chb.2024.108352

Kücking, F., Hübner, U., Przysucha, M., Hannemann, N., Kutza, J.-O., Moelleken, M. et al. 2024. Automation bias in AI-decision support: results from an empirical study. In German Medical Data Sciences 2024. IOS Press, Amsterdam, 317, 298–304.
https://doi.org/10.3233/SHTI240871

Lee, J., Park, S., Oh, S. and Ma, B. 2026. Can large language models automate the HAZOP process without human intervention? Safety Science194, 107039.
https://doi.org/10.1016/j.ssci.2025.107039

Lyell, D. and Coiera, E. 2017. Automation bias and verification complexity: a systematic review. Journal of the American Medical Informatics Association24(2), 423–431.
https://doi.org/10.1093/jamia/ocw105

Martin, H., James, J. and Chadee, A. 2025. Exploring large language model AI tools in construction project risk assessment: Chat GPT limitations in risk identification, mitigation strategies, and user experience. Journal of Construction Engineering and Management151(9), 04025119.
https://doi.org/10.1061/JCEMD4.COENG-16658

Nyqvist, R., Peltokorpi, A. and Seppänen, O. 2024. Can ChatGPT exceed humans in construction project risk management? Engineering, Construction and Architectural Management31(13), 223–243.
https://doi.org/10.1108/ECAM-08-2023-0819

Oral, M., Alboga, Ö., Aydınlı, S. and Erdis, E. 2026. Usability of large language models for building construction safety risk assessment. Engineering, Construction and Architectural Management33(9), 7021–7048.
https://doi.org/10.1108/ECAM-08-2024-1143

Parasuraman, R. and Riley, V. 1997. Humans and automation: use, misuse, disuse, abuse. Human Factors: The Journal of the Human Factors and Ergonomics Society39(2), 230–253.
https://doi.org/10.1518/001872097778543886

Qiao, H., Vermeulen, J., Fitzmaurice, G. and Matejka, J. 2025. To use or not to use: impatience and overreliance when using generative AI productivity support tools. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan. Association for Computing Machinery, 1–18.
https://doi.org/10.1145/3706598.3714103

Romeo, G. and Conti, D. 2026. Exploring automation bias in human–AI collaboration: a review and implications for explainable AI. AI & SOCIETY41, 259–278.
https://doi.org/10.1007/s00146-025-02422-7

Sun, N. and Kalar, D. 2025. Gemini at work: knowledge workers’ perceptions and assessment of productivity gains. In Proceedings of the 2025 ACM Designing Interactive Systems Conference, Funchal, Madeira, Portugal. Association for Computing Machinery, 

Back to Issue

Choose new content alerts