GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data Security

6Citations
Citations of this article
18Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Schema matching constitutes a pivotal phase in the data ingestion process for contemporary database systems. Its objective is to discern pairwise similarities between two sets of attributes, each associated with a distinct data table. This challenge emerges at the initial stages of data analytics, such as when incorporating a third-party table into existing databases to inform business insights. Given its significance in the realm of database systems, schema matching has been under investigation since the 2000s. This study revisits this foundational problem within the context of large language models. Adhering to increasingly stringent data security policies, our focus lies on the zero-shot and few-shot scenarios: the model should analyze only a minimal amount of customer data to execute the matching task, contrasting with the conventional approach of scrutinizing the entire data table. We emphasize that the zero-shot or few-shot assumption is imperative to safeguard the identity and privacy of customer data, even at the potential cost of accuracy. The capability to accurately match attributes under such stringent requirements distinguishes our work from previous literature in this domain.

Cite

CITATION STYLE

APA

Liu, X., Wang, R., Song, Y., & Kong, L. (2024). GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data Security. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 5476–5486). Association for Computing Machinery. https://doi.org/10.1145/3637528.3671602

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free