Abstract
Large multi-disciplinary scientific projects that inform government policy and have a high public profile are often exposed to high levels of scrutiny. Such projects rely on a range of input datasets and modelling software packages and generate high volumes of output data, which are presented as summarised results in published reports. Defending the scientific integrity of project reporting requires that all project results have demonstrable integrity with clear evidence of the workflows and processes used to generate them, i.e. they must implement structured data management including provenance capture and storage. Provenance data capture forms part of effective data management. The reporting of data provenance needs to occur in all workflows within a project and crucially needs support from project management, and adoption by project staff so that provenance chains are unbroken at every step, thus providing demonstrable integrity. Even when project funds and milestones are allocated to provenance tasks, such as ensuring staff store project datasets in managed locations and generate standardised dataset metadata records, data provenance capture has often been poor. This indicates that the barrier to the adoption of useful data provenance tasks is still significant. The development and application of automated systems, which capture and report provenance without additional user effort, are therefore of critical importance in helping to lower this barrier thus easing cultural change in data management. Even if a project or organisation has motivation, has made the case, established a vision, and developed plans to implement provenance management, buy-in from all project staff is still required for success. This is because provenance chains containing information about data lifecycles need to be unbroken for all results, thus requiring involvement from all project staff. Some, perhaps the majority, of project processes cannot be automated, thus they will require significant manual effort in order to be included in provenance management. This paper outlines previous best-practice regarding CSIRO's data management approach as demonstrated by the Murray Darling Basin Sustainable Yields project, and reflects on their shortcomings, such as the lack of adequate provenance capture, with improvements suggested. It then describes several automated provenance management tools that employ semantic web technologies and preserve the identity of provenance reports and datasets; which may be used to help with bottom-up practice adoption. The automated provenance management tools can provide well-defined, automated processes, which may help to lower the barriers preventing cultural change for data management at the project and organisational level. It is hoped that the improved data management practices and the automated tools discussed here can inform current and new high-profile projects, such as the Bioregional Assessments program, to attain a higher quality of demonstrable data integrity through more robust provenance management.
Author supplied keywords
Cite
CITATION STYLE
Car, N. J., Hartcher, M. G., & Stenson, M. P. (2013). Driving data management cultural change via automated provenance management systems. In Proceedings - 20th International Congress on Modelling and Simulation, MODSIM 2013 (pp. 2173–2179). Modelling and Simulation Society of Australia and New Zealand Inc. (MSSANZ). https://doi.org/10.36334/modsim.2013.k5.car
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.