Abstract
We compared the performance of large language models (LLMs) and humans with various levels of expertise in child investigative interviewing on tasks related to question formulation. Two tasks were employed: a static Interview Excerpt Task where participants (60 psychologists, 60 naive participants, GPT-4, and Llama-2) formulated follow-up questions to 100 interview excerpts, and a dynamic Avatar Interviewing Task where participants (32 professionals, 32 students, and GPT-4) conducted 10-min interviews with AI-driven child avatars. In the dynamic task, LLMs used fewer recommended questions (M = 8.69 vs. 18.75) and more non-recommended questions (M = 17.69 vs. 6.81) than professionals. Conversely, in the static task, GPT-4 outperformed psychologists, using more invitations (67.8% vs. 5.4%) and fewer option-posing questions (3.7% vs. 31.4%). While LLMs demonstrated strong question formulation skills in controlled environments, they struggled with adaptive dialogs.
Cite
CITATION STYLE
Järvilehto, L., Sun, Y., Aiba, N., Haginoya, S., Hallström, H., Korkman, J., & Santtila, P. (2026). Large Language Model (LLM) and Human Performance in Child Investigative Interviewing Question Formulation Tasks. Behavioral Sciences and the Law, 44(1), 142–163. https://doi.org/10.1002/bsl.70029
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.