Abstract
We investigate the possibility of leveraging side information for improving quality control over crowd-sourced data. We extend the GLAD model, which governs the probability of correct labeling through a logistic function in which worker expertise counteracts item difficulty, by systematically encoding different types of side information, including worker information drawn from demographics and personality traits, item information drawn from item genres and content, and contextual information drawn from worker responses and labeling sessions. Modeling side information allows for better estimation of worker expertise and item difficulty in sparse data situations and accounts for worker biases, leading to better prediction of posterior true label probabilities. We demonstrate the efficacy of the proposed framework with overall improvements in both the true label prediction and the unseen worker response prediction based on different combinations of the various types of side information across three new crowd-sourcing datasets. In addition, we show the framework exhibits potential of identifying salient side information features for predicting the correctness of responses without the need of knowing any true label information.
Cite
CITATION STYLE
Jin, Y., Carman, M., Kim, D., & Xie, L. (2017). Leveraging Side Information to Improve Label Quality Control in Crowd-Sourcing. In Proceedings of the 5th AAAI Conference on Human Computation and Crowdsourcing, HCOMP 2017 (pp. 79–88). AAAI Press. https://doi.org/10.1609/hcomp.v5i1.13315
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.