Abstract
Challenging behaviors in children with autism is a serious clinical condition, oftentimes leading to aggression or self-injurious actions. The Revised Family Observation Schedule 3 rd Edition (FOS-R-III) is an intensive and fine-grained scale used to observe and analyze the behaviors of individuals with autism, which facilitates the diagnosis and monitoring of autism severity. Previous AI-based approaches for automated behavior analysis in autism often focused on predicting facial expressions and body movements without generating a clinically meaningful scale, mostly utilizing visual information. In this study, we propose a deep-learning based algorithm with audio-visual multimodal-data clinically coded with the FOS-R-III, named AV-FOS model. Our proposed AV-FOS model leverages transformer-based structure and self-supervised learning to intelligently recognize Interaction Styles (IS) in the FOS-R-III scale from subjects' video recordings. This enables the automatic generation of the FOS-R-III measures with clinically acceptable accuracy. We explore the IS recognition using a multimodal large language model, GPT4V, with prompt engineering provided with FOS-R-III measure definitions as the baseline for this study and compare with other vision-based deep learning algorithms. We believe this research represents a significant advancement in autism research and clinical accessibility. The proposed AV-FOS and our FOS-R-III dataset will serve as a gateway toward the digital health era for future AI models related to autism.
Author supplied keywords
Cite
CITATION STYLE
Zhao, Z., Chung, E., Chung, K. M., & Park, C. H. (2025). AV-FOS: Transformer-Based Audio-Visual Multimodal Interaction Style Recognition for Children With Autism Using the Revised Family Observation Schedule 3rd Edition (FOS-R-III). IEEE Journal of Biomedical and Health Informatics, 29(9), 6238–6250. https://doi.org/10.1109/JBHI.2025.3542066
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.