Abstract
Functional Hazard Assessment (FHA) is a cornerstone of certification-oriented safety assessment, yet it remains predominantly manual, expert-dependent, and time-intensive. This paper evaluates the extent to which a Large Language Model (LLM) can support early-phase FHA generation for the Guided Precision Airdrop System (GPADS) using a controlled pipeline that integrates Retrieval-Augmented Generation (RAG) and schema-based guardrails to enforce FHA structure and traceability. A multi-dimensional evaluation framework compares an expert-developed reference FHA against an LLM-generated FHA in terms of functional coverage, failure-mode representation, hazard granularity, severity assignment, blind-spot contributions, and documentation quality. At a function level similarity threshold of 0.7 - defined as a textual similarity score computed over GPADS function labels and functional descriptions using a fuzzy string-matching approach based on Levenshtein distance - the model recovered all expert-defined functions (100% recall) with 80% precision, indicating substantial functional reconstruction while introducing additional functions beyond the expert baseline. The LLM produced a more fine-grained hazard set (48 items versus 19 in the reference) but exhibited systematic severity inflation and reduced failure-mode diversity, reinforcing the need for expert calibration. A second qualitative study shows that expert review can consolidate and correct the raw output, yielding an intermediate hazard set (35 items) with severity distributions re-aligned to certification logic, while retaining the efficiency benefits of automated drafting. Overall, the results provide focused single-case evidence that a human-in-the-loop LLM workflow can substantially accelerate exploratory hazard elicitation and structured FHA documentation, while experts remain essential for scope control, consolidation, and severity assessment.
Author supplied keywords
Cite
CITATION STYLE
Bourguignon, V. H. O., Pleffken, D. R., Rocha, G. C., & Garcia, J. S. D. (2026). Human-in-the-loop large language models for functional hazard assessment: A controlled RAG-based pipeline and evidence from a GPADS case study. Results in Engineering, 30. https://doi.org/10.1016/j.rineng.2026.111313
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.