OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation

3Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Online Public Opinion Reports consolidate news and social media for timely crisis management by governments and enterprises. While large language models (LLMs) enable automated report generation, this specific domain lacks formal task definitions and corresponding benchmarks. To bridge this gap, we define the Automated Online Public Opinion Report Generation (OPOR-Gen) task and construct OPOR-Bench, an event-centric dataset with 463 crisis events across 108 countries (comprising 8.8 K news articles and 185 K tweets). To evaluate report quality, we propose OPOR-Eval, a novel agent-based framework that simulates human expert evaluation. Validation experiments show OPOR-Eval achieves a high Spearman’s correlation (ρ = 0.70) with human judgments, though challenges in temporal reasoning persist. This work establishes an initial foundation for advancing automated public opinion reporting research.

Cite

CITATION STYLE

APA

Yu, J., Xu, Y., Li, H., Li, J., Zhu, L., Shen, H., & Shi, L. (2026). OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation. Computers, Materials and Continua, 87(1). https://doi.org/10.32604/cmc.2025.073771

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free