Abstract
In a modernized statistical production process, non-traditional data sources such as ‘big' data are increasingly being considered either as the main or as a supplementary source for official statistics. Their use however has brought new sorts of challenges: messy datasets, duplicate entries, missing information and misspellings to name a few. In many cases, there is also no unique identifier which can be used to unambiguously identify a record for the purpose of data integration. These challenges can be compounded by a non-English based alphabet like the Persian/Farsi alphabet used in Iran. In this paper, two innovative methods have been elaborated to address such data challenges. More specifically, the application of probabilistic record linkage using an ACSII coding system is an innovative way to deal with both data challenges and lack of unique identifier simultaneously. Moreover, text mining is an innovative way to address categorization and grouping systems that are not suitable for statistical purposes. Both innovative approaches can improve the accuracy and coherency of datasets and for data integration result in higher quality datasets. Results of research undertaken by the authors show the innovations lead to more effective data integration and improve the quality of the resulting official statistics. The innovations have wide applicability especially in non-English alphabet countries.
Author supplied keywords
Cite
CITATION STYLE
Fayyaz, S., & Hadizadeh, R. (2020). Innovations from Iran: Resolving quality issues in the integration of administrative and big data in official statistics. Statistical Journal of the IAOS, 36(4), 1015–1030. https://doi.org/10.3233/SJI-200756
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.