Abstract
This study examines whether machine learning can be used to predict player performance in the Czech National Basketball League (NBL) using only information available before a game. Publicly available player- and team-level data were collected from the official NBL website and FIBA LiveStats, and the target variable was defined as Performance Index Rating (PIR), a widely used indicator of overall player efficiency in European basketball. To ensure practical usability and methodological validity, all predictor variables were constructed exclusively from pre-game information, including lagged and rolling statistics from previous matches, and explicit steps were taken to prevent data leakage. Three supervised learning models were compared: Random Forest, Extreme Gradient Boosting (XGBoost), and a simple neural network. Model performance was evaluated using R², MAE, RMSE, and MSE. Among the tested models, XGBoost achieved the best overall performance, although the predictive accuracy remained moderate rather than high, indicating that pre-game prediction of PIR is a challenging task in a smaller European league context. Feature-importance and SHAP analyses showed that prior PIR, recent form, minutes-related indicators, and selected matchup variables contributed most strongly to predictions. The study contributes to the growing literature on sports analytics by providing a transparent and reproducible pre-game prediction pipeline for Czech basketball, and the article-specific code, analytical notebook, and merged player-game dataset are made publicly available in an online repository.
Author supplied keywords
Cite
CITATION STYLE
Salaj, Š. (2026). Machine Learning Prediction of Player Performance in the Czech National Basketball League. Studia Sportiva, 20(1). https://doi.org/10.5817/StS2026-1-17
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.