Abstract
Predicting the affinity profiles of nucleic acid-binding proteins directly from the protein sequence is a challenging problem. We present a statistical approach for learning the recognition code of a family of transcription factors or RNA-binding proteins (RBPs) from high-throughput binding data. Our method, called affinity regression, trains on protein binding microarray (PBM) or RNAcompete data to learn an interaction model between proteins and nucleic acids using only protein domain and probe sequences as inputs. When trained on mouse homeodomain PBM profiles, our model correctly identifies residues that confer DNA-binding specificity and accurately predicts binding motifs for an independent set of divergent homeodomains. Similarly, when trained on RNAcompete profiles for diverse RBPs, our model correctly predicts the binding affinities of held-out proteins and identifies key RNA-binding residues, despite the high level of sequence divergence across RBPs. We expect that the method will be broadly applicable to modeling and predicting paired macromolecular interactions in settings where high-throughput affinity data are available.
Cite
CITATION STYLE
Pelossof, R., Singh, I., Yang, J. L., Weirauch, M. T., Hughes, T. R., & Leslie, C. S. (2015). Affinity regression predicts the recognition code of nucleic acid-binding proteins. Nature Biotechnology, 33(12), 1242–1249. https://doi.org/10.1038/nbt.3343
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.