FICSIM: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction

1Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

As language models become capable of processing increasingly long and complex texts, there has been growing interest in their application within computational literary studies. However, evaluating the usefulness of these models for such tasks remains challenging due to the cost of fine-grained annotation for long-form texts and the data contamination concerns inherent in using public-domain literature. Current embedding similarity datasets are not suitable for evaluating literary-domain tasks because of a focus on coarse-grained similarity and primarily on very short text. We assemble and release FICSIM, a dataset, of long-form, recently written fiction, including scores along 12 axes of similarity informed by author-produced metadata and validated by digital humanities scholars. We evaluate a suite of embedding models on this task, demonstrating a tendency across models to focus on surface-level features over semantic categories that would be useful for computational literary studies tasks. Throughout our data-collection process, we prioritize author agency and rely on continual, informed author consent.

Cite

CITATION STYLE

APA

Johnson, N., Bertsch, A., Deal, M. E., & Strubell, E. (2025). FICSIM: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction. In EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2025 (pp. 25228–25246). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-emnlp.1375

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free