Abstract
Highlights: What are the main findings? Machine learning models trained on embedding features derived from the Google AlphaEarth Foundations dataset demonstrated strong predictive capability for forest biomass, particularly in broad-leaved and coniferous forests, as validated through five-fold cross-validation. These models successfully captured large-scale spatial patterns of forest biomass distribution. Spatially predicational biomass maps revealed clear landscape-scale gradients, with lower biomass values predominantly occurring in fragmented forest patches and forest edges near urbanized areas, while higher biomass levels were concentrated within continuous, intact forest regions that are relatively distant from human disturbance. What are the implications of the main findings? Embedding-based remote sensing models offer an effective and efficient framework for monitoring biomass dynamics in Yunhe Forestry Station, especially in regions where field surveys or the acquisition of multi-source remote sensing data are constrained by terrain accessibility or logistical limitations. By leveraging embeddings that integrate information from diverse Earth observation sensors, this study demonstrates an optional and scalable methodology for forest biomass estimation, highlighting the potential of representation learning to advance large-area forest carbon assessment and management. Spatial predictions of forest biomass at regional scale in forests are critical to evaluate the effects of management practices across environmental gradients. Although multi-source remote sensing combined with machine learning has been widely applied to estimate forest biomass, these approaches often rely on complex data acquisition and processing workflows that limit their scalability for large-area assessments. To improve the efficiency, this study evaluates the potential of annual multi-sensor satellite embeddings derived from the AlphaEarth Foundations model for forest biomass prediction. Using field inventory data from 89 forest plots at the Yunhe Forestry Station in Zhejiang Province, China, we assessed and compared the performance of four machine learning algorithms: Random Forest (RF), Support Vector Regression (SVR), Multi-Layer Perceptron Neural Networks (MLPNN), and Gaussian Process Regression (GPR). Model evaluation was conducted using repeated 5-fold cross-validation. The results show that SVR achieved the highest predictive accuracy in broad-leaved and mixed forests, whereas RF performed best in coniferous forests. When all forest types were modeled together, predictive performance was consistently limited across algorithms, indicating substantial heterogeneity (e.g., structure, environment, and topography) among forest types. Spatial prediction maps across Yunhe Forestry Station revealed ecologically coherent patterns, with higher biomass values concentrated in intact forests with less human disturbance and lower biomass primarily occurring in fragmented forests and near urban regions. Overall, this study highlights the potential of embedding-based remote sensing for regional forest biomass estimation and suggests its utility for large-scale forest monitoring and management.
Author supplied keywords
Cite
CITATION STYLE
Jin, C., Jiang, X., Wen, L., Wu, C., Xu, X., & Jiao, J. (2026). Assessing the Utility of Satellite Embedding Features for Biomass Prediction in Subtropical Forests with Machine Learning. Remote Sensing, 18(3). https://doi.org/10.3390/rs18030436
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.