Amino acid sequence encodes protein abundance shaped by protein stability at reduced synthesis cost
Citations Over TimeTop 18% of 2024 papers
Abstract
Understanding what drives protein abundance is essential to biology, medicine, and biotechnology. Driven by evolutionary selection, an amino acid sequence is tailored to meet the required abundance of a proteome, underscoring the intricate relationship between sequence and functional demand. Yet, the specific role of amino acid sequences in determining proteome abundance remains elusive. Here we show that the amino acid sequence alone encodes over half of protein abundance variation across all domains of life, ranging from bacteria to mouse and human. With an attempt to go beyond predictions, we trained a manageable-size Transformer model to interpret latent factors predictive of protein abundances. Intuitively, the model's attention focused on the protein's structural features linked to stability and metabolic costs related to protein synthesis. To probe these relationships, we introduce MGEM (Mutation Guided by an Embedded Manifold), a methodology for guiding protein abundance through sequence modifications. We find that mutations which increase predicted abundance have significantly altered protein polarity and hydrophobicity, underscoring a connection between protein structural features and abundance. Through molecular dynamics simulations we revealed that abundance-enhancing mutations possibly contribute to protein thermostability by increasing rigidity, which occurs at a lower synthesis cost.
Related Papers
- → The primary structure of a plant storage protein: zein(1981)164 cited
- → “De-novo” amino acid sequence elucidation of protein G′e by combined “Top-Down” and “Bottom-Up” mass spectrometry(2015)14 cited
- → Nucleotide sequence of the Mn-stabilizing protein gene of the thermophilic cyanobacterium Synechococcus elongatus(1993)15 cited
- → The complete cDNA coding sequence for the mouse CDEI binding protein(1993)17 cited
- → Identification of peptides within a known protein sequence using COMSEQ analysis of data containing multiple sequences(1991)2 cited