Abstract
Multi-modal face anti-spoofing (FAS) aims to detect genuine human presence by extracting discriminative liveness cues from multiple modalities, such as RGB, infrared (IR), and depth images, to enhance the robustness of biometric authentication systems. However, because data from different modalities are typically captured by various camera sensors and under diverse environmental conditions, multi-modal FAS often exhibits significantly larger distribution discrepancies across training and testing domains compared to single-modal FAS. Furthermore, during the inference stage, multi-modal FAS confronts even greater challenges when one or more modalities are unavailable or inaccessible. To address these issues, we propose a Cross-modal Transition-guided Network (CTNet) for robust multi-modal FAS. Our motivation stems from that, within a single modality, live faces exhibit smaller visual variations than spoof faces, and cross-modal feature transitions are more consistent for live samples than for spoof ones. Upon this insight, we propose learning consistent cross-modal feature transitions among live samples to construct a generalized feature space. Next, we introduce learning inconsistent cross-modal transitions between live and spoof samples to effectively detect out-of-distribution (OOD) attacks during inference. To further address the issue of missing modalities, we propose learning complementary IR and depth features from the RGB modality as auxiliary modalities. Extensive experiments demonstrate that the proposed CTNet outperforms previous multi-modal FAS methods across most protocols.
Author supplied keywords
Cite
CITATION STYLE
Chong, J. X., Hsu, F. Y., Hsu, M. T., Lin, Y. T., Chien, K. H., Hsu, C. T., & Huang, P. K. (2026). Multi-modal face anti-spoofing via cross-modal feature transitions. Expert Systems with Applications, 310. https://doi.org/10.1016/j.eswa.2026.131292
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.