Abstract
The DNA sequences of an organism play an important influence on its transcription and translation process, thus affecting its protein production and growth rate. Due to the complexity of DNA, it was extremely difficult to predict the macroscopic characteristics of organisms. However, with the rapid development of machine learning in recent years, it becomes possible to use powerful machine learning algorithms to process and analyze biological data. Based on the synthetic DNA sequences of a specific microbe, E. coli, I designed a process to predict its protein production and growth rate. By observing the properties of a data set constructed by previous work, I chose to use supervised learning regressors with encoded DNA sequences as input features to perform the predictions. After comparing different encoders and algorithms, I selected three encoders to encode the DNA sequences as inputs and trained seven different regressors to predict the outputs. The hyper-parameters are optimized for three regressors which have the best potential prediction performance. Finally, I successfully predicted the protein production and growth rates, with the best R2 score 0.55 and 0.77, respectively, by using encoders to catch the potential features from the DNA sequences.
Cite
CITATION STYLE
Zhao, S. (2021). Prediction of Protein Expression and Growth Rates by Supervised Machine Learning. Natural Science, 13(08), 301–330. https://doi.org/10.4236/ns.2021.138025
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.