Abstract
Generalization of models to out-of-distribution (OOD) data has captured tremendous attention recently. Specifically, compositional generalization, i.e., whether a model generalizes to new structures built of components observed during training, has sparked substantial interest. In this work, we investigate compositional generalization in semantic parsing, a natural test-bed for compositional generalization, as output programs are constructed from sub-components. We analyze a wide variety of models and propose multiple extensions to the attention module of the semantic parser, aiming to improve compositional generalization. We find that the following factors improve compositional generalization: (a) using contextual representations, such as ELMO and BERT, (b) informing the decoder what input tokens have previously been attended to, (c) training the decoder attention to agree with pre-computed token alignments, and (d) downsampling examples corresponding to frequent program templates. While we substantially reduce the gap between in-distribution and OOD generalization, performance on OOD compositions is still substantially lower.
Cite
CITATION STYLE
Oren, I., Herzig, J., Gupta, N., Gardner, M., & Berant, J. (2020). Improving compositional generalization in semantic parsing. In Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020 (pp. 2482–2495). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.findings-emnlp.225
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.