Abstract
Constructing a universal moral code for artificial intelligence (AI) is challenging because human cultures have different values, norms, and social practices. We therefore argue that AI systems should adapt to culture based on observation: Just as a child raised in a particular culture learns the specific values, norms, and behaviors of that culture, we propose that an AI system operating in a particular human community could similarly learn them as well. How AI systems might accomplish this from observing and interacting with humans has remained an open question. Here, we propose using inverse reinforcement learning (IRL) as a method for AI agents to acquire culturally relevant values implicitly from humans. We test our approach using an experimental paradigm in which AI agents use IRL to learn different reward functions, which govern the agents’ actions, by learning from variations in the altruistic behavior of human subjects from two cultural groups in an online game requiring real-time decision making. We show that an AI agent learning from a particular human cultural group can acquire the altruistic characteristics reflective of that group’s average behavior, and can generalize to new scenarios requiring altruistic judgments. Our results provide a proof-of-concept demonstration that AI agents can be endowed with the ability to learn culturally-typical behaviors and values directly from observing human behavior.
Cite
CITATION STYLE
Oliveira, N., Li, J., Khalvati, K., Barragan, R. C., Reinecke, K., Meltzoff, A. N., & Rao, R. P. N. (2025). Culturally-attuned AI: Implicit learning of altruistic cultural values through inverse reinforcement learning. PLOS ONE, 20(12 December). https://doi.org/10.1371/journal.pone.0337914
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.