Abstract
Noun phrases and Relation phrases in open knowledge graphs are not canonicalized, leading to an explosion of redundant and ambiguous subject-relation-object triples. Existing approaches to solve this problem take a two-step approach. First, they generate embedding representations for both noun and relation phrases, then a clustering algorithm is used to group them using the embeddings as features. In this work, we propose Canonicalizing Using Variational Autoencoders (CUVA), a joint model to learn both embeddings and cluster assignments in an end-to-end approach, which leads to a better vector representation for the noun and relation phrases. Our evaluation over multiple benchmarks shows that CUVA outperforms the existing state-of-the-art approaches. Moreover, we introduce CANONICNELL, a novel dataset to evaluate entity canonicalization systems.
Cite
CITATION STYLE
Dash, S., Rossiello, G., Bagchi, S., Mihindukulasooriya, N., & Gliozzo, A. (2021). Open Knowledge Graphs Canonicalization using Variational Autoencoders. In EMNLP 2021 - 2021 Conference on Empirical Methods in Natural Language Processing, Proceedings (pp. 10379–10394). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2021.emnlp-main.811
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.