Toward Explainable Facial Expression Recognition Using Face Action Units: XFER-AU

1Citations
Citations of this article
4Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Face expression recognition (FER) has been extensively explored by the research community, with comparisons between different FER models typically relying on accuracy metrics. To make model decisions more interpretable, evaluation can be complemented with the usage of (model-agnostic) explainability tools, leading to explainable FER models (XFER). However, the effectiveness of explainability tools in the context of XFER is seldom evaluated. This paper proposes a framework, entitled “face expression recognition explainability evaluation framework, based on action units” (XFER-AU), for evaluating the quality of explanations provided by different explainability tools in the context of XFER. To establish a comparison term, XFER-AU automatically generates a FER explanation ground truth, based on facial action units (AU). The proposed framework thus enables the comparison of explainability tools in terms of their ability to explain the decision made by an FER model. To perform the comparison, an evaluation metric called weighted explainability score (WES) is proposed, which takes into account the number of ground truth AUs covered and the precision of the explainability map produced. The proposed framework is used to compare several explainability tools, notably Local Interpretable Model-agnostic Explanation (LIME), Gradient-weighted Class Activation Mapping (Grad-CAM), SHapley Additive exPlanation (SHAP), Average Removal/Aggregation (AVG), Minus and PLUS explanation (MinPLUS) and Randomised Input Sampling for Explanation (RISE). The proposed XFER-AU results report that, for the tested datasets, LIME explanations tend to better align with the ground truth. In more challenging scenarios, the FER models struggle to effectively point to the facial AUs relevant to explaining the observed expression, thereby reducing the model’s decision confidence. Results also show that even when an expression is correctly classified, older FER models, such as those based on VGG-16, support their decision on a smaller number of the relevant AUs when compared to more recent models, such as EfficientNet-B0.

Cite

CITATION STYLE

APA

Verlekar, T. T., Goyal, A., & Lobato Correia, P. (2026). Toward Explainable Facial Expression Recognition Using Face Action Units: XFER-AU. IEEE Access, 14, 3625–3638. https://doi.org/10.1109/ACCESS.2025.3649062

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free