Abstract
The construction industry persistently underperforms in hazard recognition, often leading to severe workplace injuries due to unrecognized hazards. With the recent advancements in Artificial Intelligence (AI) and the emergence of Large Language Models (LLM), the construction sector has begun exploring these technologies for various applications. However, a systematic comparison of popular LLMs to evaluate their effectiveness in identifying construction hazards remains unexplored. Additionally, previous studies have primarily focused on assessing LLMs using textual input and output, leaving their performance with visual inputs underexplored. This study addresses this gap by systematically assessing and comparing the hazard recognition performance of five widely used LLMs using construction case images. The findings establish a baseline standard for LLMs in construction hazard identification through zero-shot learning and reveal that LLMs do not perform significantly well in this context. Additionally, the study provides valuable insights into the reliability and potential applications of LLMs for enhancing hazard recognition in the construction industry.
Author supplied keywords
Cite
CITATION STYLE
Chaudhary, N., Uddin, S. M. J., Tamanna, M., Albert, A., & Shahid, A. R. B. (2025). How Reliable Are Large Language Models? Zero-Shot Detection of Construction Hazards. In Proceedings of the International Symposium on Automation and Robotics in Construction (pp. 649–656). International Association for Automation and Robotics in Construction (IAARC). https://doi.org/10.22260/ISARC2025/0085
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.