Abstract
The core driver of enterprise operations is data, making data lineage crucial for data management. It not only tracks data flow but also links data sources, workflows, applications, and decision-making, improving efficiency and governance. However, current data lineage parsing methods face challenges like high costs, long development cycles, and poor generalization, especially for non-SQL scripts. In this paper, we introduce an innovative approach leveraging pre-trained large language models (LLMs) to overcome these bottlenecks in data lineage parsing. LLMs are employed across the entire parsing pipeline, encompassing prompt construction, lineage extraction, and result standardization. Specifically, this study developed a few-shot prompting method incorporating error cases to optimize parsing performance across various types of scripts. Additionally, a collaborative Chain of Thought (CoT) and multi-expert prompting framework was designed to further enhance parsing accuracy at the operator level. The proposed approach was empirically validated using LLMs of different parameter scales on datasets comprising multiple script types (SQL, Python, Shell, Flume, etc.). The experimental results show that LLMs with 10 billion and 100 billion parameters achieved over 95% accuracy in table-level lineage parsing when utilizing the newly designed prompts. Furthermore, 100-billion-parameter LLMs exhibited substantial accuracy improvements at the operator level. Our method reinforces the feasibility and practicality in advancing data lineage parsing methodologies.
Author supplied keywords
Cite
CITATION STYLE
Li, Z., Guo, W., Gao, Y., Yang, D., & Kang, L. (2025). A Large Language Model-Based Approach for Data Lineage Parsing. Electronics (Switzerland), 14(9). https://doi.org/10.3390/electronics14091762
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.