Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?

1Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Charts are ubiquitous as they help people understand and reason with data. Recently, various downstream tasks, such as chart question answering, chart captioning, etc. have emerged. Large Vision-Language Models (LVLMs) show promise in tackling these tasks, but their qualitative evaluation is costly and time-consuming, limiting real-world deployment. While using LVLMs as judges to assess chart comprehension capabilities of other LVLMs could streamline evaluation processes, challenges like proprietary datasets, restricted access to powerful models, and evaluation costs hinder their adoption in industrial settings. To this end, we present a comprehensive evaluation of 13 open-source LVLMs as judges for diverse chart comprehension and reasoning tasks. We design both pairwise and pointwise evaluation tasks covering criteria like factual correctness, informativeness, and relevancy. Additionally, we analyze LVLM judges based on format adherence, positional consistency, length bias, and instruction-following. We focus on cost-effective LVLMs (≤ 9B parameters) suitable for both research and commercial use, following a standardized evaluation protocol and rubric to measure the LVLM judge accuracy. Experimental results reveal notable variability: while some open LVLM judges achieve GPT-4level evaluation performance (about 80% agreement with GPT-4 judgments), others struggle (below 10% agreement). Our findings highlight that state-of-the-art open-source LVLMs can serve as cost-effective automatic evaluators for chart-related tasks, though biases such as positional preference and length bias persist.

Cite

CITATION STYLE

APA

Laskar, M. T. R., Islam, M. S., Mahbub, R., Masry, A., Rahman, M., Bhuiyan, M. A. H., … Huang, J. X. (2025). Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 6, pp. 1203–1216). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.acl-industry.83

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free