Abstract
Recently, the risk of information disclosure is increasing significantly. Accordingly, privacy-preserving data mining (PPDM) is being actively studied to obtain accurate mining resultswhile preserving the data privacy.We here focus on secure similar document detection (SSDD), which identifies similar documents of two parties when each party does not disclose its own sensitive documents to the another party. In this paper, we propose an efficient two-step protocol that exploits a feature selection as a lower-dimensional transformation, and we present discriminative feature selections to maximize the performance of the protocol. The proposed protocol consists of two steps: the filtering step and the postprocessing step. For the feature selection, we first consider the simplest one, random projection (RP), and propose its two-step solution, SSDD-RP.We then present two discriminative feature selections and their solutions: SSDD-LF which selects a few dimensions locally frequent in the current querying vector and SSDD-GF which selects ones globally frequent in the set of all document vectors.We finally propose a hybrid one, SSDD-HF, which takes advantage of both SSDD-LF and SSDD-GF.We empirically showthat the proposed two-step protocol significantly outperforms the previous one-step protocol by three or four orders of magnitude.
Cite
CITATION STYLE
Kim, S. P., Gil, M. S., Kim, H., Choi, M. J., Moon, Y. S., & Won, H. S. (2017). Efficient two-step protocol and its discriminative feature selections in secure similar document detection. Security and Communication Networks, 2017. https://doi.org/10.1155/2017/6841216
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.