How Reliable Are Large Language Models? Zero-Shot Detection of Construction Hazards

0Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

Abstract

The construction industry persistently underperforms in hazard recognition, often leading to severe workplace injuries due to unrecognized hazards. With the recent advancements in Artificial Intelligence (AI) and the emergence of Large Language Models (LLM), the construction sector has begun exploring these technologies for various applications. However, a systematic comparison of popular LLMs to evaluate their effectiveness in identifying construction hazards remains unexplored. Additionally, previous studies have primarily focused on assessing LLMs using textual input and output, leaving their performance with visual inputs underexplored. This study addresses this gap by systematically assessing and comparing the hazard recognition performance of five widely used LLMs using construction case images. The findings establish a baseline standard for LLMs in construction hazard identification through zero-shot learning and reveal that LLMs do not perform significantly well in this context. Additionally, the study provides valuable insights into the reliability and potential applications of LLMs for enhancing hazard recognition in the construction industry.

Cite

CITATION STYLE

APA

Chaudhary, N., Uddin, S. M. J., Tamanna, M., Albert, A., & Shahid, A. R. B. (2025). How Reliable Are Large Language Models? Zero-Shot Detection of Construction Hazards. In Proceedings of the International Symposium on Automation and Robotics in Construction (pp. 649–656). International Association for Automation and Robotics in Construction (IAARC). https://doi.org/10.22260/ISARC2025/0085

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free