Large language models (LLMs) are increasingly adopted in occupational health and safety (OHS) to accelerate workplace risk assessment, yet their integration raises a dilemma between efficiency and verification. The study examines this trade-off in an LLM-assisted workflow using data from 111 student assignment files from an educational setting, analyzed at the session level, yielding 96‒121 analytic sessions depending on measure availability. Efficiency was measured via perceived time savings versus an estimated manual baseline, while the verification-burden proxy was operationalized as an AI‒human disagreement in quantitative scoring (probability and impact on 1‒5 scales). User acceptance was captured via the utility and satisfaction / intention to reuse Likert scales.
The results show substantial perceived efficiency gains: the median LLM-assisted time was 20 minutes versus the median manual time of 120 minutes (median saving 0.84). The verification-burden proxy remained non-trivial: median disagreement was 0.50, and 26.8% of sessions exhibited disagreement ≥1.0. Acceptance remained high (mean satisfaction 4.06/5). Notably, 40.6% of the sessions showed a possible overreliance pattern, with high satisfaction coexisting with high disagreement. In a robust ordinary least squares (OLS) regression, utility was the strongest predictor of satisfaction, while disagreement reduced satisfaction; time saving was not significant.
Perceived speed gains may thus coexist with hidden assurance costs and possible trust miscalibration, highlighting the need for verification structures and targeted training for sustainable digital safety governance.
Aladağ, H. 2023. Assessing the accuracy of ChatGPT use for risk management in construction projects. Sustainability, 15(22), 16071.
https://doi.org/10.3390/su152216071
Choudhury, A. and Shamszare, H. 2024. The impact of performance expectancy, workload, risk, and satisfaction on trust in ChatGPT: cross-sectional survey analysis. JMIR Human Factors, 11(1), e55399.
https://doi.org/10.2196/55399
Collier, Z. A., Gruss, R. J. and Abrahams, A. S. 2024. How good are large language models at product risk assessment? Risk Analysis, 45(4), 766–789.
https://doi.org/10.1111/risa.14351
Falegnami, A., Tomassi, A., Corbelli, G., Nucci, F. S. and Romano, E. 2024. A generative artificial-intelligence-based workbench to test new methodologies in organisational health and safety. Applied Sciences, 14(24), 11586.
https://doi.org/10.3390/app142411586
Goddard, K., Roudsari, A. and Wyatt, J. C. 2012. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127.
https://doi.org/10.1136/amiajnl-2011-000089
Klingbeil, A., Grützner, C. and Schreck, P. 2024. Trust and reliance on AI ‒ an experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior, 160, 108352.
https://doi.org/10.1016/j.chb.2024.108352
Kücking, F., Hübner, U., Przysucha, M., Hannemann, N., Kutza, J.-O., Moelleken, M. et al. 2024. Automation bias in AI-decision support: results from an empirical study. In German Medical Data Sciences 2024. IOS Press, Amsterdam, 317, 298–304.
https://doi.org/10.3233/SHTI240871
Lee, J., Park, S., Oh, S. and Ma, B. 2026. Can large language models automate the HAZOP process without human intervention? Safety Science, 194, 107039.
https://doi.org/10.1016/j.ssci.2025.107039
Lyell, D. and Coiera, E. 2017. Automation bias and verification complexity: a systematic review. Journal of the American Medical Informatics Association, 24(2), 423–431.
https://doi.org/10.1093/jamia/ocw105
Martin, H., James, J. and Chadee, A. 2025. Exploring large language model AI tools in construction project risk assessment: Chat GPT limitations in risk identification, mitigation strategies, and user experience. Journal of Construction Engineering and Management, 151(9), 04025119.
https://doi.org/10.1061/JCEMD4.COENG-16658
Nyqvist, R., Peltokorpi, A. and Seppänen, O. 2024. Can ChatGPT exceed humans in construction project risk management? Engineering, Construction and Architectural Management, 31(13), 223–243.
https://doi.org/10.1108/ECAM-08-2023-0819
Oral, M., Alboga, Ö., Aydınlı, S. and Erdis, E. 2026. Usability of large language models for building construction safety risk assessment. Engineering, Construction and Architectural Management, 33(9), 7021–7048.
https://doi.org/10.1108/ECAM-08-2024-1143
Parasuraman, R. and Riley, V. 1997. Humans and automation: use, misuse, disuse, abuse. Human Factors: The Journal of the Human Factors and Ergonomics Society, 39(2), 230–253.
https://doi.org/10.1518/001872097778543886
Qiao, H., Vermeulen, J., Fitzmaurice, G. and Matejka, J. 2025. To use or not to use: impatience and overreliance when using generative AI productivity support tools. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan. Association for Computing Machinery, 1–18.
https://doi.org/10.1145/3706598.3714103
Romeo, G. and Conti, D. 2026. Exploring automation bias in human–AI collaboration: a review and implications for explainable AI. AI & SOCIETY, 41, 259–278.
https://doi.org/10.1007/s00146-025-02422-7
Sun, N. and Kalar, D. 2025. Gemini at work: knowledge workers’ perceptions and assessment of productivity gains. In Proceedings of the 2025 ACM Designing Interactive Systems Conference, Funchal, Madeira, Portugal. Association for Computing Machinery,