Comparison of optimized Markov Decision Process using Dynamic Programming and Temporal Differencing - A reinforcement learning approach

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Reinforcement learning is one of the promising approaches for operations research problems. The incoming inspection process in any manufacturing plant aims to control quality, reduce manufacturing costs, eliminate scrap, and process failure downtimes due to nonconforming raw materials. Prediction of the raw material acceptance rate can regulate the raw material supplier selection and improve the manufacturing process by filtering out non-conformities. This paper presents a Markov model developed to estimate the probability of the raw material being accepted or rejected in an incoming inspection environment. The proposed forecasting model is further optimized for efficiency using the two reinforcement learning algorithms (dynamic programming and temporal differencing). The results of the two optimized models are compared, and the findings are discussed.

Cite

CITATION STYLE

APA

Mani, A., Bakar, S. A., Krishnan, P., & Yaacob, S. (2021). Comparison of optimized Markov Decision Process using Dynamic Programming and Temporal Differencing - A reinforcement learning approach. In Journal of Physics: Conference Series (Vol. 2107). IOP Publishing Ltd. https://doi.org/10.1088/1742-6596/2107/1/012026

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free