Reinforcement Learning–Based Building Control Considering Decision-Makers′ Preferences Between Energy Use and Indoor Environmental Quality

0Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Most reinforcement learning (RL)–based building control studies design multiple reward functions with relative weights to balance conflicting objectives, such as energy efficiency and indoor environmental quality. However, existing research often relies on arbitrarily assigned weights determined by RL model developers. This practice limits the ability to systematically and objectively reflect the preferences of building decision-makers, including owners, facility managers, administrators, and residents. The objective of this study is to develop a systematic method to reflect the preferences of building decision-makers in RL-based control. To achieve this objective, the Analytic Hierarchy Process (AHP) was employed to quantify user preferences, and K-means clustering was applied to categorize them into distinct preference groups. The resulting preference types are used to assign reward weights in a double deep Q-network (DDQN) algorithm for optimal indoor climate control. To validate the proposed approach, preference data were collected from decision-makers in a university dormitory, including the administrative manager, facility manager, and residents. Using AHP, the relative importance of energy performance, thermal comfort, and indoor air quality (IAQ) was quantified. K-means clustering was then applied to identify five representative preference types: energy performance conscious, thermal comfort—IAQ conscious, IAQ conscious, thermal comfort conscious, and energy performance—thermal comfort conscious. The centroid of each cluster was used to assign reward weights to the corresponding DDQN models. The performance of the RL models was evaluated in a testbed chamber simulating a dormitory environment. Compared with the baseline model using uniform reward weights (PMV: 0.39, total energy consumption: 3.78 kWh), the energy performance conscious model minimized energy consumption to 0.27 kWh while maintaining a warmer indoor thermal environment (PMV: 1.88). In contrast, the thermal comfort conscious model achieved a highly comfortable environment (PMV: 0.12), but with significantly higher energy use (6.40 kWh). These results demonstrate that the proposed hybrid approach can objectively quantify and classify decision-maker preferences and effectively incorporate them into RL-based building control strategies.

Cite

CITATION STYLE

APA

Kim, S. H., & Moon, H. J. (2026). Reinforcement Learning–Based Building Control Considering Decision-Makers′ Preferences Between Energy Use and Indoor Environmental Quality. Indoor Air, 2026(1). https://doi.org/10.1155/ina/4881534

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free