On the relation of causality- versus correlation-based feature selection on model fairness

20Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

As machine learning models are used increasingly in the educational domain, ensuring that they are fair and do not discriminate against certain groups or individuals is imperative. Although there are a few recent attempts to ensure fairness in these models, the majority of fairness literature tends to overlook the feature selection (FS) process despite its critical role as one of the foundational steps in the machine learning pipeline. Moreover, traditional FS methods identify features by examining the correlational relationships between predictive features and the target variable without seeking to uncover causal connections between them. To address these issues, we compare for four openly available datasets - -two educational ones and two benchmark datasets regularly used in the fairness literature - -the impact of these two different ways of FS (i.e., causality- versus correlation-based) on the performance and fairness of the resulting models. Our results show that causality-based FS generally leads to fairer models, while the models built after correlation-based FS manifest higher performance.

Cite

CITATION STYLE

APA

Saarela, M. (2024). On the relation of causality- versus correlation-based feature selection on model fairness. In Proceedings of the ACM Symposium on Applied Computing (pp. 56–64). Association for Computing Machinery. https://doi.org/10.1145/3605098.3636018

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free