Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning

7Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Offline safe reinforcement learning (OSRL) involves learning a decision-making policy to maximize rewards from a fixed batch of training data to satisfy pre-defined safety constraints. However, adapting to varying safety constraints during deployment without retraining remains an under-explored challenge. To address this challenge, we introduce constraint-adaptive policy switching (CAPS), a wrapper framework around existing offline RL algorithms. During training, CAPS uses offline data to learn multiple policies with a shared representation that optimize different reward and cost trade-offs. During testing, CAPS switches between those policies by selecting at each state the policy that maximizes future rewards among those that satisfy the current cost constraint. Our experiments on 38 tasks from the DSRL benchmark demonstrate that CAPS consistently outperforms existing methods, establishing a strong wrapper-based baseline for OSRL.

Cite

CITATION STYLE

APA

Chemingui, Y., Deshwal, A., Wei, H., Fern, A., & Doppa, J. (2025). Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 39, pp. 15722–15730). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v39i15.33726

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free