Abstract
The growing capacity to handle vast amounts of data, combined with a shift in service delivery models, has improved scalability and efficiency in data analytics, particularly in multi-tenant environments. Data are treated as digital products and processed through orchestrated service-based data pipelines. However, advancements in data analytics do not find a counterpart in data governance techniques, leaving a gap in the effective management of data throughout the pipeline lifecycle. This gap highlights the need for innovative service-based data pipeline management solutions that prioritize balancing data quality and data protection. The framework proposed in this paper optimizes service selection and composition within service-based data pipelines to maximize data quality while ensuring compliance with data protection requirements, expressed as access control policies. Given the NP-hard nature of the problem, a sliding-window heuristic is defined and evaluated against the exhaustive approach and a baseline modeling the state of the art. Our results demonstrate a significant reduction in computational overhead, while maintaining high data quality.
Author supplied keywords
Cite
CITATION STYLE
Polimeno, A., Braghin, C., Anisetti, M., & Ardagna, C. A. (2025). Maximizing data quality while ensuring data protection in service-based data pipelines. Journal of Big Data, 12(1). https://doi.org/10.1186/s40537-025-01118-5
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.