Abstract
This paper provides a comprehensive and descriptive machine learning model for forecasting the yield of biochar, bio-oil, and pyrolysis gases in biomass-plastic co-pyrolysis systems. Three algorithms, namely, Linear Regression (LR), Decision Tree (DT), and Extreme Gradient Boosting (XGBoost), were comparatively explored using compositional, catalytic, and operational variables as predictors. Training-Test validation, Monte Carlo cross-validation (20 repeated 80/20 splits), and uncertainty analysis were used to assess the model's robustness and generalizability. LR showed poor predictive ability and low test R2 values, suggesting it fails to capture nonlinear thermochemical interactions. DT showed moderate results but exhibited mild overfitting. Conversely, XGBoost performed superiorly in terms of predictive accuracy with high test R2 of 0.995 (biochar), 0.984 (pyrolysis oil), and 0.987 (pyrolysis gas), and had a small RMSE (0.608 to 2.208) and consistent Monte Carlo performance. Partial dependence plots, permutation importance, and SHAP bee swarm analysis were used to explain the results. The explainability analysis showed that blend ratio primarily controlled biochar yield, whereas temperature, ash content, and fixed carbon were the key drivers of bio-oil and gas production. The proposed framework combines high predictive accuracy with mechanistic interpretability, providing a trustworthy decision-support tool for optimal sustainable co-pyrolysis processes.
Author supplied keywords
Cite
CITATION STYLE
Nguyen, D., Chen, W. H., Guerrero-Pérez, M. O., Rodríguez-Castellón, E., Nguyen, V. Q., Islam, A., … Hoang, A. T. (2026). Explainable and parsimonious machine learning models for predicting product yield from biomass–plastic co-pyrolysis. Applied Thermal Engineering, 302. https://doi.org/10.1016/j.applthermaleng.2026.131703
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.