CausalLink: An Interactive Evaluation Framework for Causal Reasoning

0Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We present CausalLink, an innovative evaluation framework that interactively assesses the causal reasoning skill to identify the correct intervention in conversational language models. Each CausalLink test case creates a hypothetical environment in which the language models are instructed to apply interventions to entities whose interactions follow predefined causal relations generated from controllable causal graphs. Our evaluation framework isolates causal capabilities from the confounding effects of world knowledge and semantic cues. We evaluate a series of LLMs in a scenario featuring movements of geometric shapes and discover that models start to exhibit reliable reasoning on two or three variables at the 14-billion-parameter scale. However, the performance of state-of-the-art models such as GPT4o degrades below random chance as the number of variables increases. We identify and analyze several key failure modes.

Cite

CITATION STYLE

APA

Feng, J., & Rudzicz, F. (2025). CausalLink: An Interactive Evaluation Framework for Causal Reasoning. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 22313–22326). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.1147

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free