Abstract
In application domains requiring mission-critical decision-making, such as finance and robotics, the optimal policy derived by reinforcement learning (RL) often hinges on a preference for risk management. Yet, the dynamic nature of risk measures poses considerable challenges to achieving generalization and adaptation of risk-sensitive policies in the context of RL. In this paper, we propose a risk-conditioned RL framework that enables rapid policy adaptation to varying risk measures via a unified risk representation, Weighted Value-at-Risk (WV@R). To sample risk measures that avoid undue optimism, we construct a risk proposal network employing a conditional adversarial auto-encoder and a normalizing flow. This network establishes coherent representations for risk measures, ensuring the monotonicity of the quantile representations of risk measures. Through experiments with locomotion, finance, and self-driving scenarios, we show that our framework is capable of adapting to a range of risk measures, achieving comparable performance to the baselines individually trained for each measure. The framework often outperforms the baselines, especially in the cases when exploration is required during training but risk-aversion is favored during evaluation.
Cite
CITATION STYLE
Yoo, G., Park, J., & Woo, H. (2024). Risk-Conditioned Reinforcement Learning: A Generalized Approach for Adapting to Varying Risk Measures. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, pp. 16513–16521). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v38i15.29589
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.