Human-in-the-loop large language models for functional hazard assessment: A controlled RAG-based pipeline and evidence from a GPADS case study

1Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Functional Hazard Assessment (FHA) is a cornerstone of certification-oriented safety assessment, yet it remains predominantly manual, expert-dependent, and time-intensive. This paper evaluates the extent to which a Large Language Model (LLM) can support early-phase FHA generation for the Guided Precision Airdrop System (GPADS) using a controlled pipeline that integrates Retrieval-Augmented Generation (RAG) and schema-based guardrails to enforce FHA structure and traceability. A multi-dimensional evaluation framework compares an expert-developed reference FHA against an LLM-generated FHA in terms of functional coverage, failure-mode representation, hazard granularity, severity assignment, blind-spot contributions, and documentation quality. At a function level similarity threshold of 0.7 - defined as a textual similarity score computed over GPADS function labels and functional descriptions using a fuzzy string-matching approach based on Levenshtein distance - the model recovered all expert-defined functions (100% recall) with 80% precision, indicating substantial functional reconstruction while introducing additional functions beyond the expert baseline. The LLM produced a more fine-grained hazard set (48 items versus 19 in the reference) but exhibited systematic severity inflation and reduced failure-mode diversity, reinforcing the need for expert calibration. A second qualitative study shows that expert review can consolidate and correct the raw output, yielding an intermediate hazard set (35 items) with severity distributions re-aligned to certification logic, while retaining the efficiency benefits of automated drafting. Overall, the results provide focused single-case evidence that a human-in-the-loop LLM workflow can substantially accelerate exploratory hazard elicitation and structured FHA documentation, while experts remain essential for scope control, consolidation, and severity assessment.

Cite

CITATION STYLE

APA

Bourguignon, V. H. O., Pleffken, D. R., Rocha, G. C., & Garcia, J. S. D. (2026). Human-in-the-loop large language models for functional hazard assessment: A controlled RAG-based pipeline and evidence from a GPADS case study. Results in Engineering, 30. https://doi.org/10.1016/j.rineng.2026.111313

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free