An empirical study of validating synthetic data for formula generation

0Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Large language models (LLMs) can be leveraged to help write formulas in spreadsheets, but formula data resources are scarce, impacting both the base performance of pre-trained models and limiting the ability to fine-tune them. Given a corpus of formulas, we can use another model to generate synthetic natural language utterances for fine-tuning. However, it is important to validate whether the natural language (NL) generated by the LLM is accurate for it to be beneficial for fine-tuning. In this paper, we provide empirical results on the impact of validating these synthetic training examples with surrogate objectives that evaluate the accuracy of the synthetic annotations. We demonstrate that validation improves performance over raw data across four models (2 open and 2 closed weight). Interestingly, we show that although validation tends to prune more challenging examples, it increases the complexity of problems that models can solve after being fine-tuned on validated data.

Cite

CITATION STYLE

APA

Singh, U., Kanade, A., Cambronero, J., Khatry, A., Gulwani, S., Le, V., … Verbruggen, G. (2025). An empirical study of validating synthetic data for formula generation. In 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025 (pp. 7062–7069). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-naacl.391

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free