Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation

1Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Given the growing influence of language model-based agents on high-stakes societal decisions, from public policy to healthcare, ensuring their beneficial impact requires understanding the far-reaching implications of their suggestions. We propose a proof-of-concept framework that projects how model-generated advice could propagate through societal systems on a macroscopic scale over time, enabling more robust alignment. To assess the long-term safety awareness of language models, we also introduce a dataset of 100 indirect harm scenarios, testing models' ability to foresee adverse, non-obvious outcomes from seemingly harmless user prompts. Our approach achieves not only over 20% improvement on the new dataset but also an average win rate exceeding 70% against strong baselines on existing safety benchmarks (AdvBench, SafeRLHF, WildGuardMix), suggesting a promising direction for safer agents.

Cite

CITATION STYLE

APA

Sun, C., Zhang, D., Zhai, C. X., & Ji, H. (2025). Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 6422–6434). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.332

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free