Abstract
The growing demand for biopharmaceutical products reflects their effectiveness in medical treatments. However, developing new biopharmaceuticals remains a major bottleneck, often taking up to a decade before market approval. Machine learning (ML) models have the potential to accelerate this process, but their success depends on access to large and diverse data sets for training. Multi-fidelity ML techniques offer a promising solution by integrating abundant, low-cost, and less accurate low-fidelity (LF) data with limited, expensive, and more accurate high-fidelity (HF) data. In this framework, LF data capture global system trends, while HF data refine and align model predictions with the available ground truth. Such integration can substantially reduce development costs and timelines by minimizing the need to acquire HF data, for example, through extensive experimental campaigns. This work reviews developments in surrogate modeling within the biopharmaceutical context, including Gaussian processes, neural networks, and physics-informed approaches. It also provides practical recommendations for identifying appropriate LF and HF data. Existing research has primarily focused on upstream processing and drug discovery, highlighting opportunities to extend these methods to other stages, like downstream processing. While Gaussian processes and neural networks remain the most frequently used models, emerging architectures such as Transformer and diffusion models present promising directions for future research.
Author supplied keywords
Cite
CITATION STYLE
Golzarijalal, M., Aickelin, U., & Otte, E. (2026, June 1). Addressing Small Data Challenges in Biopharmaceutical Development and Manufacturing: A Mini Review of Multi-Fidelity Techniques. Biotechnology and Bioengineering. John Wiley and Sons Inc. https://doi.org/10.1002/bit.70213
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.