Parallel Rule Discovery from Large Datasets by Sampling

18Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Rule discovery from large datasets is often prohibitively costly. The problem becomes more staggering when the rules are collectively defined across multiple tables. To scale with large datasets, this paper proposes a multi-round sampling strategy for rule discovery. We consider entity enhancing rules (REEs) for collective entity resolution and conflict resolution, which may carry constant patterns and machine learning predicates. We sample large datasets with accuracy bounds a and B such that at least a% of rules discovered from samples are guaranteed to hold on the entire dataset (i.e., precision), and at least B% of rules on the entire dataset can be mined from the samples (i.e., recall). We also quantify the connection between support and confidence of the rules on samples and their counterparts on the entire dataset. To scale with the number of tuple variables in collective rules, we adopt deep Q-learning to select semantically relevant predicates. To improve the recall, we develop a tableau method to recover constant patterns from the dataset. We parallelize the algorithm such that it guarantees to reduce runtime when more processors are used. Using real-life and synthetic data, we empirically verify that the method speeds up REE discovery by 12.2 times with sample ratio 10% and recall 82%.

Author supplied keywords

Cite

CITATION STYLE

APA

Fan, W., Han, Z., Wang, Y., & Xie, M. (2022). Parallel Rule Discovery from Large Datasets by Sampling. In Proceedings of the ACM SIGMOD International Conference on Management of Data (pp. 384–398). Association for Computing Machinery. https://doi.org/10.1145/3514221.3526165

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free