Abstract
The increasing adoption of cloud infrastructures and multitenancy in big data environments has reshaped the scalability and efficiency of data analytics. However, multitenancy introduces new challenges in data governance, particularly regarding privacy and security. Existing research often treats data quality and security as separate concerns, focusing either on improving the accuracy and reliability of data or on protecting sensitive information. This article offers a thorough data governance system that takes a holistic approach to addressing these difficulties for contemporary data pipelines. By including data protection and functional requirements into the service pipeline, our method allows for the selection and assembly of services that maximize data quality while maintaining privacy and security standards. We provide a parametric heuristic as a solution to the NP-hard service selection issue, guaranteeing maximal data retention while maintaining regulatory compliance. The framework is evaluated through experiments on a real-world dataset, demonstrating its ability to balance data quality and protection. This work contributes to the ongoing efforts to enhance privacy-aware data governance in multitenant cloud-based analytics ecosystems.
Cite
CITATION STYLE
Chatterjee, S. (2023). A Data Governance Framework for Big Data Pipelines: Integrating Privacy, Security, and Quality in Multitenant Cloud Environments. Technix International Journal for Engineering Research, 10(5). https://doi.org/10.56975/tijer.v10i5.158181
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.