Full text
Benefit and cost of a protein Manish Kushwaha1, Wolfram Liebermeister2, Elad Noor2, Kirill Sechkar4, Diana Széliová5 1Université Paris-Saclay, INRAE, AgroParisTech, Micalis Institute, 78352 Jouy-en-Josas, France 2Université Paris-Saclay, INRAE, MaIAGE, 78350 Jouy-en-Josas, France 3Department of Plant and Environmental Sciences, Weizmann Institute of Science, 76100 Rehovot, Israel 4Department of Engineering Science, University of Oxford, Parks Road, Oxford OX1 3PJ, UK 5Department of Analytical Chemistry, University of Vienna, Austria Abstract In this chapter we discuss the cellular growth rate effects of expressing a protein. We show how protein costs and benefits can be defined and how they reflect both the protein’s molecule properties and cell-wide changes in resource allocation. For enzymes, the benefit depends on enzyme efficiencies, while protein cost depends mostly on protein size and amino acid composition, representing energy, and material needed to produce the protein, possibly conceptualized as opportunity costs. After an overview of protein production and the necessary resources, we study the optimal expression of a metabolic enzyme by a simple cost-benefit model. Then we discuss how empirical costs and benefits can be defined and measured for the Lac operon in Escherichia coli. Finally, we discuss predictions from constraint-based and mechanistic cell models. In an appendix, we give details about experimental techniques for measuring transcription and translation, controlling expression levels, and measuring growth rates. Keywords: protein expression - cost-benefit optimization - evolution in the lab - growth rate measurement - resource allocation Contributions: Discussed with Martin Lercher, Hugo Dourado, Ohad Golan. Early version reviewed by Hugo Dourado and Ohad Golan. Michela Pauletti drew most of the figures in this chapter. Figure 11.1 was redrawn from [1] with kind permission by the authors. Version August 2025 was reviewed by Maud Hofmann, Ohad Golan, and Andreas Kremling. To cite this chapter: M. Kushwaha, W. Liebermeister, E. Noor, K. Sechkar, and D. Széliová. Benefit and cost of a protein (Version January 2026). doi: 10.5281/zenodo.17955350. Chapter from: The Economic Cell Collective (2026). Economic Principles in Cell Biology. No commercial publisher | Online open access book | doi: 10.5281/zenodo.8156386 The authors are listed in alphabetical order. This is a chapter from the open textbook “Economic Principles in Cell Biology”. Free download from principlescellphysiology.org/book-economic-principles/. Lecture slides for this chapter are available on the website. c 2026 The Economic Cell Collective. Licensed under Creative Commons License CC-BY-SA 4.0. An online open access book. No publisher has been paid. doi: 10.5281/zenodo.8156386
1 Chapter overview ◦How are global resource allocation and cell fitness reflected in local choices such as determining the expression level of a single protein? Separating between the cost and benefit of a specific protein’s expression helps describe this local optimization quantitatively. ◦Effective costs and benefits of a single protein may be defined operationally, as measured increases or decreases in cell growth. For enzymes, the factors that determine cost and benefit include enzyme efficiency and protein cost (size, energy, amino acids, ribosomes needed to make proteins, as well as all kinds of opportunity costs). ◦Marginal cost (or benefit) is defined as the derivative of the cost (or benefit) curve, i.e. the effect of small changes in expression. At the point of maximal fitness, the marginal cost and marginal benefit must be equal. ◦We describe a simple cost-benefit model where benefit is defined as the production flux and cost as the enzyme level. Based on the same idea, we then present empirical definitions and measurements of the costs and benefits of the lac-operon in Escherichia coli. ◦We discuss predictions from constraint-based and mechanistic cell models and compare them to the empirical definitions. ◦In an appendix, we describe methods for measuring transcription, translation, and growth rates, as well as for controlling expression to observe costs/benefits for different protein levels. 11.1. Effects of protein expression on cell growth 11.1.1. How do changes in a protein level affect cell growth? Protein production in cells is a complex process that involves transcription, translation, modifications, and transport, and it depends on various other processes. When protein expression levels are changing, this will affect ribosome demand and metabolic production fluxes. Proteins also have forward effects, not only by performing various cellular functions but also by competing for space among each other and with other compounds. Changes in protein levels may have a variety of indirect effects, including rearrangements of gene expression, which depend on regulatory mechanisms and will impact the cell-wide allocation of resources. In models, these changes may either be described physically–as a result of regulatory mechanisms–or based on optimality principles. Our models from previous chapters are aimed at predicting such rearrangements. Having learned how to model resource allocation in cells, we now ask a simple question: When a cell expresses a heterologous protein or overexpresses (or underexpresses) a native protein, how will this change its growth rate? Since cell growth depends on many cellular processes, the answer to this question can be quite complex. However, according to simple economic logic, each protein under each given condition should exhibit an optimal expression level. If a protein level is too low, the protein cannot sufficiently exert its function, and if it is too high, the protein consumes too much of the cellular resources, which will then be lacking for processes elsewhere. As often in life, the optimum is a compromise, a “sweet spot” at which cells reach a maximal growth rate (or optimize some other relevant fitness objective). At this optimal protein level, any increase or decrease would result in a decrease in the growth rate. Of course, we may wonder: will cells actually achieve this optimal protein level? We can test this in experiments. Figure 11.1 shows the expression of a number of proteins in different bacterial species. In this set of experiments, the amount of each selected protein was varied by placing its coding gene under the control of an inducible promoter, and cell growth rates were measured for several expression levels. All measured protein/growth curves were negatively curved (i.e. concave), where the maximum growth rate was found at an intermediate protein expression level. Moreover, the wild-type level of each protein–that is, the level in natural, non-modified cells–was close to its optimal level. By setting the enzyme level to this optimal value, wild-type cells often come close to the maximal possible growth rate. Figure 11.2 shows this for one specific protein, ATP synthase
2 (A) (B) 0 1 2 3 0 0.5 1 Titrated level / WT level µtitrated / µWT E. coli, citrate synthase Glucose/Acetate Acetate 012345 0 0.5 1 Titrated level / WT level µtitrated / µWT E. coli, ATP synthase Glucose Succinate (C) (D) 0 1 2 3 0 0.5 1 Titrated level / WT level µtitrated / µWT S. typhimurium, IIAGLC Maltose 012345 0 0.5 1 Titrated level / WT level µtitrated / µWT L. lactis LDH PFK LAS GAPDH Figure 11.1: Growth rate effects of protein expression in different bacterial species – The panels show different examples of optimal enzyme expression. (A) Citrate synthase in Escherichia coli. (B) ATP synthase in E. coli (C) PTS transporter system (glucose-specific subunit IIA) in Salmonella typhimurium. (D) Several glycolytic enzymes in Lactococcus lactis. Figure redrawn from [1] with kind permission by the authors. The ATP synthase data stem from [2]. In all cases, “Titrated levels” refer to proteins under the control of an IPTG-inducible promoter. in E. coli, under a large range of growth conditions. The close match between the observed and maximal growth rates suggests that the regulatory mechanisms behind this protein have evolved to adjust the protein levels to support maximal growth. And this protein is not an exception. Wild-type levels of other proteins have been found to be growth-optimal as well, including those of some efflux pumps [3] and the enzyme methionine synthase [4]. Of course, this principle–wild-type protein levels are the protein levels that maximize growth–does not hold for all proteins and all types of cells. In yeast cells, for example, some protein levels were found to be growth-optimal while others were not, and this also depends on the growth conditions [5,6]. But given the complex functioning of real cells, it still remains interesting to see what would be the optimal protein levels in theory, and by what general logic we can understand them. The growth rate effects of proteins are related to various relevant questions: 1. Understanding protein usage and protein levels in naturally evolved cells. What is the best expression level for each protein in wild-type cells (e.g. the titration levels that maximize the functions in Figure 11.1)? And what are the conditions under which each single protein should be expressed? 2. Understanding selection pressures on protein levels, which shape evolution. How do deviations from the optimal protein levels affect the fitness objective? For example, in Figure 11.1, how strongly do the curves decrease when deviating from the optimal point? Knowing the curvature, we can get an idea about the cost of fluctuations and, therefore, flexibility in protein expression.
3 0 0.2 0.4 0.6 0.811.2 0 0.2 0.4 0.6 0.8 1 1.2 alanine arabinose cytosine galactose glucosamine glucose glutamine glycine LB maltose mannose ornithine pyruvate ribose sorbitol succinate sucrose acetate 2-ketoglutarate arginine asparagine fructose glucose-6P glycerol lactate mannitol trehalose Growth rate, wild-type [hr−1] Max. growth rate, titrated strain [hr−1] Figure 11.2: Wild-type abundances of ATP synthase in E. coli lead to near-optimal growth rates – The plot compares the growth rate of wild-type cells in different environments to the corresponding maximal growth rate achievable by varying the abundance of ATP synthase. Only in four of the 27 growth conditions shown, the wild-type growth rate of E. coli deviates by more than 10% from the maximal growth rate. Figure redrawn from [1] with kind permission by the authors. The ATP synthase data were taken from [2]. 3. Expressing heterologous pathways in cells in biotechnology. Optimal choices of protein levels are also important for biotechnology. Optimizing protein levels for maximal production of valuable compounds or for maximal growth rates (e.g., by expression via controllable gene promoters) is a relevant biotechnological task and an example of cellular economics. Whenever proteins are expressed artificially, it is important to understand – and possibly limit – the effects on cell growth. A large growth deficit will slow down the reproduction of the cells and, therefore, total production. Moreover, a lower growth rate may lead to evolutionary selection against the engineered strain: mutants with disabled heterologous genes may experience lower growth defects and therefore take over the population by outcompeting the engineered cells. But the question remains: How do these protein/growth curves arise? What determines their shape, and how do they differ between proteins? In this chapter, we consider the growth rate effects of proteins in detail. Following the economic perspective of this book, our main focus is not on how protein levels are regulated in cells but on how they contribute to the cell’s functioning. That is, we do not inquire about the mechanistic causes but about the incentives for increasing or decreasing a protein level. There are several reasons for looking at cells in this way. First, if there is an evolutionary selection for high growth rates, the resulting regulatory systems will support optimal "expression programs". Second, even if we don’t fully understand how regulation systems function, we may use optimality principles as a way out, hoping that they can serve as a good approximation. 11.1.2. Protein cost and benefit The curves in Figure 11.1 may look simple, but they may reflect complex rearrangements in the entire cell. Protein production is a complex process that involves transcription, translation, modifications, and transport and depends on various other processes. When protein levels change, all these processes may be adapted, the demand for ribosomes changes, and metabolic production fluxes may be rerouted. Proteins typically serve a specific function, and changing their abundance perturbs this function, too. Moreover, there are various indirect effects; for example, due to the competition for cellular space, which is crowded with other proteins, membranes, polymers, and small molecules (see
4 (A) (B) (C) futile overexpression external regulator growth rate μ ∆µ ∆e marginal growth rate ∆µ/∆e≈∂µ/∂e < 0 expression e growth rate µ optimal growth marginal growth rate ∂µ/∂e = 0 expression e growth rate µ Figure 11.3: Quantification of protein/growth rate effects – (A) An externally controlled expression of a protein (with expression level e) influences the cell growth rate µ. (B) Forced expression of an idle protein decreases the growth rate. In wild-type cells under a selection for growth, such a protein should not be expressed. At small expression levels, the growth rate decreases approximately linearly (first-order effect). (C) In the case of a beneficial protein, the growth rate becomes maximal at some intermediate expression level, and small variations around this level make the growth rate decrease. Near the optimal point, the curve can be approximated by a quadratic curve (second-order effect). In the optimal point the curve has a zero slope, and small increases or decreases of the protein level would leave the growth rate almost unchanged (compare Figure 11.1). Section 2.6.1 in [7]). Mechanistically, these cell-wide rearrangements depend on regulatory mechanisms and impact the allocation of resources. Accordingly, there are two different ways to describe this reallocation in models: either physically – that is, as resulting from regulation mechanisms – or based on optimality principles. Our models from previous chapters can be used to predict such rearrangements. One way to describe the effects of protein levels on cell growth is to assume that each protein has a benefit and a cost. On the one hand, we assume that a protein contributes to the functioning of the cell and that its presence has a positive effect on growth. On the other hand, protein production and maintenance require resources, which puts a burden on cell growth. These resources include precursors, energy, and labor time of the biosynthesis machinery. Furthermore, packing the cellular volume with more proteins would decrease the diffusion rate of small molecules in the cell (slowing down metabolism as a whole), which means that the space proteins occupy is yet another limited resource. The more abundant a protein is in a cell, the less space and material remains for other molecules that also perform important functions. Below, we assume that any protein in a cell has a cost and that protein expression must be justified by a benefit, or the protein will not be expressed. At the optimal expression level, cost and benefit must be “balanced” such that any small changes in protein levels are fitness-neutral; that is, their additional costs and benefits cancel out. In principle, such costs and benefits can be studied by controlling the abundance of a protein via an external regulator, for example a compound that can activate the protein’s gene promoter, and measuring the resulting cell growth rate (Fig. 11.3(A). An “idle” protein, which provides no benefit to the cell in the current conditions, will only contribute a cost. The higher its expression level, the more slowly the cell will grow, and the optimum growth rate is reached when the protein is not expressed (Fig. 11.3(B)). In the case of a “useful” protein, there is also a benefit; when the expression level becomes very small, the benefit will show a sharp drop, with leads to an optimum at some intermediate expression level (Fig. 11.3(A)). Following this idea, and using such experimental data, we may split a protein’s growth effects into separate cost and benefit terms, hoping to understand each of these terms through simple considerations about the enzyme’s physical properties such as its molecular mass, size, or enzyme efficiency. Notions of protein costs and benefits can be defined in different ways, depending on their purpose:
5 1. Empirical definition and measurements. A common operational definition is based on experiments in which the expression level of a protein is externally controlled, and the resulting changes in cell growth rate are measured. By comparing experiments in which beneficial and costly effects are varied, one can try to disentangle them to better understand what causes them in the cell. 2. Model simulations. Instead of real experiments, we may also perform computer experiments based on models. The models described in the previous chapters 7 in [7] and 10 in [7] allow us to simulate the change in enzyme levels and compute the resulting achievable growth rate. 3. Theoretical definition and calculation. Aside from numerical calculations, mathematical models can help us understand general principles (see box “Qualities of a model” in 5 in [7]). If we made a model and know what assumptions went into it, we can understand it more easily than a real cell. Some model properties can be analyzed without simulations: metabolic control analysis, for example, can reveal general principles of optimal enzyme allocation and how it should be adapted upon an external change. In all three cases, the aim is to split cell-wide fitness effects into simple cost and benefit contributions attributed to a single protein. To justify this splitting, we often need to assume that protein changes are small compared to the entire proteome. This allows us, in many cases, to model protein costs and benefits using linear or quadratic approximations. In Section 11.2 of this chapter, we consider cell growth as an example objective and describe how the influence of protein expression can be measured. You can find more details about experimental methods in the appendix. In Section 11.3, we introduce a simple cost/benefit model for a single enzyme; we then apply it to the Lac operon in E. coli, for which protein costs and benefits have been measured (Section 11.4). Finally, we discuss some general principles relating empirical costs and benefits to (direct or indirect) mechanistic causes, show how protein/growth curves can be predicted using the models discussed earlier in the book (Section 11.5), and ask how protein expression is regulated mechanistically and how this may lead to growth-optimal states (Section 11.6). 11.2. Protein levels and growth rate: mechanisms and observations 11.2.1. Protein production and the necessary resources To understand how protein levels affect cell fitness, we first need to know the steps in protein production and the necessary resources. The amino acid sequence information for making a protein is (usually) encoded in DNA. DNA is transcribed to messenger RNA (mRNA) and then translated to protein. These processes require three basic ingredients: precursors, molecular machines that assemble these precursors, and energy to power these processes. The precursors for transcription are nucleotides, and for translation amino acids. These can be synthesized by the cell itself or be taken up from the medium. Transcription and translation are catalyzed by large molecular machines, such as RNA polymerases and ribosomes. Finally, these machines require energy in the form of ATP or GDP that the cell must produce (Figure 11.4). Overall, protein translation is the most costly process. It consumes about four ATP molecules per amino acid, whereas transcription uses about two ATP molecules per nucleotide (BioNumbers ID 105375). Additionally, the catalysts involved in translation are larger; in bacteria, ribosomes are 2700 kDa (BioNumbers ID 100118), compared to polymerases with 480 kDa (Bionumbers ID 104925). Furthermore, an mRNA molecule is used to make multiple proteins. Due to these cost differences, the translation machinery occupies a larger portion of the cell’s proteome than transcription machinery (roughly 5-10 times larger in E. coli, see proteomaps.net [8]). All of these processes contribute to the overall cost of protein production, with synthesis costs primarily determined by protein size. However, synthesis alone does not account for the full energetic burden. Proteins often require additional steps such as chaperone-assisted folding, post-translational modifications, or transport to specific cellular compartments, each of which consumes extra energy. In some cases, proteins can also impose toxic effects that further
6 benefits? additional costs? limited space molecular machines precursors and energy overhead costs transcription translation NT AA ATP ATP Figure 11.4: Steps in making a protein (see also Box 2.A in [7] in Chapter 2 in [7]) – 1. Transcription: DNA is transcribed to mRNA by RNA polymerase, which requires nucleotides (NT) and energy (typically ATP or GTP). 2. Translation: mRNA is translated into a protein by the ribosome, using amino acids (AA) and energy. These processes are constrained by the available space and require additional indirect costs, as the molecular machines catalyzing these processes need to be synthesized by the same processes, necessitating even more machines. After the protein is made, it can be beneficial for the cell or cause additional costs (e.g. costs of folding and degradation of misfolded proteins, membrane stress, by-product accumulation), depending on the type of protein and environmental conditions. For more background about biological machines and enzymes, see Chapter 2 in [7]. increase cellular energy demands. For instance, microbial production of biofuels can disrupt membrane fluidity and permeability, leading to the leakage of essential molecules such as ATP and ions [9]. Additionally, we need to consider that to make more protein, a cell also needs more molecular machines, which, in turn, requires even more molecular machines to synthesize themselves. However, the space in the cell is limited. In fact, as we have seen in Chapter 2 in [7], the protein content per cell is quite constant across conditions, but the relative composition changes. That means that to make more of one protein, the cell needs to reduce the amount of other (potentially beneficial) proteins – which can reduce growth. The limited space does not only apply to the cell as a whole but also to cellular substructures. For example, the abundance of membrane proteins is limited by the membrane area, and the abundance of mitochondrial enzymes is limited by the mitochondrial volume. 11.2.2. Cell-wide effects of protein expression: costs and benefits via different routes To understand how a single protein can affect cell growth, we need to trace its effects on other processes in the cell. A change in a single protein level may imply – that is, either cause or require – rearrangements in the entire proteome. If the total protein amount in a cell were constant, then increasing the level of one protein would cause the levels of other proteins to decrease, with secondary effects elsewhere in the cell. In contrast, if the total protein amount can vary, an increased protein level may lead to an increase in the total protein amount, which entails adjustments in the protein production machinery (polymerases, ribosomes, available charged tRNA). This, in turn, has cell-wide effects. In both cases, if a change in protein levels affects the growth rate, this will entail various other adjustments (e.g. changes in ribosome numbers, with all kinds of secondary effects). To describe the cost of a protein, including all these effects as “baggage costs” or opportunity costs of the protein, we usually assume that these costs are comparable between proteins and we can approximate them by empirical or heuristic protein cost functions. To see how protein expression affects cell growth, we need to consider the different ways in which a protein – by its presence or changes in abundance – influences the rest of the cell. If we consider a cell as a whole and zoom in on a single protein species, we notice a number of direct connections between the protein and other cell processes.
7 Experimental methods 11.A Challenges in quantifying protein expression To quantify the fitness effects of a protein in real cells, one may enforce changes in protein levels and plot these levels against the resulting changes in growth rate. However, in experiments we cannot control protein concentrations directly. To vary it indirectly, we can modulate gene copy number, transcription, or translation, for example, by titrating an inducer (e.g., IPTG with lac promoter) or using promoters with different strengths. The inducer concentration or promoter activity is then used as the input variable for cost and benefit calculations [2]. However, these measurable variables may not correspond exactly to protein concentrations because multiple processes contribute to the final protein levels. These processes can become saturated or affected by regulation [10], and these effects are gene-specific [11]. There are also cases in which gene expression correlates with inducer concentration at the population level, but individual cells show “all-or-none” expression. This means that instead of each cell gradually increasing expression, only the proportion of fully induced cells rises [12], so the cell models used in this chapter would not even apply. So how can we make sure that our calculations are done with accurate protein amounts? One solution is to measure the absolute level of our protein, for example with quantitative mass spectrometry. However, these methods are technically challenging – the measurement error are large and vary between proteins [10]. An alternative is to measure relative protein expression, for example using gel electrophoresis as done by Dekel and Alon [13] (see Section 11.4). These “direct routes of action” concern (1) How the protein is made, including the need for resources such as amino acids, energy carried by ATP, molecular machines (including ribosomes and chaperones), and mRNA templates. In eukaryotes, transporting proteins to the right place in the cell may require transporters and larger infrastructures, such as the endoplasmic reticulum. (2) The space that the protein occupies (thus reducing the available space for other cell components). (3) Running costs, for example, ATP consumption in the case of ion transporters or flagella proteins. (4) The protein’s direct function, for instance, is the catalysis of a metabolic reaction. It is through these connections that our protein influences other parts of the cell. One way to predict how these influences play out when it comes to cell growth is to use detailed whole-cell models. But how can we understand these effects intuitively, and how can we relate them to cost and benefit terms? A changing protein level will change all the direct demands (for example, ribosome demand and space demand). To describe these demands, we now make a simplifying assumption. While different proteins require resources in different amounts, e.g. due to their different molecule sizes, the way in which these resource demands affect the cell (e.g. by higher ensuing ribosome demands) do not depend on the type of protein. If this is true, then knowing the cost factors of one protein will also tell us about the cost factors of other proteins (see Box ??). Hence, the costs of proteins are assumed to depend only on general factors such as protein size, composition, lifetime, or localization. On the benefit side, unlike protein costs, protein functions differ largely between different proteins (for example, among the thousands of proteins in a cell, only one may be able to catalyze some specific reaction) – with important consequences for the entire cell! In practice, it is common to summarize all cost factors of a protein into one cost term associated with the required resources and proportional to the protein’s abundance, and to describe the protein’s function by one protein-specific benefit term. 11.3. Cost-benefit balance for an enzyme level If expressing a protein causes costs and benefits in the cell, then how can we find the “sweet spot”, the expression level where fitness is maximized? And how will this optimal value change when external conditions are changing? When external conditions are changing and make a protein more beneficial, it stands to reason that the cell should have more of it. But why exactly? And how are greater advantage and greater quantity related? A way to think
8 (A) enzyme flux benefitfitness effect enzyme cost (B) (C) (D) e benefit B cost H e fitness F=B−H e marginal benefit ∂B ∂e marginal cost ∂H ∂e Figure 11.5: Optimal enzyme level in a metabolic reaction – (A) The enzyme influences the pathway flux via its catalyzing reaction. All other enzyme levels are fixed. In an optimality problem, we define cell fitness as a difference between a benefit proportional to the pathway flux and a cost proportional to the enzyme level. Internal metabolites are shown in yellow, external metabolites are shown in orange. (B) Cost (red) and benefit (blue) as functions of the enzyme level. The point of maximal fitness (benefit minus cost) is marked by an orange star. (C) The fitness, as a function of the enzyme level, has an optimum point in which the curve has a vanishing slope. (D) The derivatives of cost and benefit functions, as functions of the enzyme level, are called marginal cost and benefit. In the optimal point, marginal cost and benefit must be equal. We can already see this in (B), where the slopes are equal in the optimal point. about this is "marginal economics", that is, considering the costs and benefits of small changes around a given state of interest. The optimal point, where any change decreases the fitness, is also a point in which very small changes have no fitness effect: since the slope of the curve is zero, moving a bit to the left or right will lead to almost no change in the growth rate. We can also see this from the protein/growth curves in Figure 11.1. How is this related to costs and benefits? If the fitness function is benefit-cost difference, then its slope (called "marginal fitness") is the difference between the slopes of the benefit and cost curves, called marginal benefit and cost. In the optimal point, where the slope of the fitness function must be zero, marginal cost and benefit cancel out: this is what we mean by saying that "cost and benefit are in balance" (see Box ??). In this logic, the “sweet spot” is a point where an infinitesimal change would have equal cost and and benefit effects, so it this point, there is no incentive to increase or decrease the enzyme level. 11.3.1. Cost-benefit balance of a metabolic enzyme – a simple additive model To see a concrete example, let us consider a metabolic pathway containing an adjustable enzyme as shown in Fig. 11.5 (A) and assume that the enzyme level is optimized [14]. We further assume that the pathway contributes to cell fitness in two ways, described by separate fitness terms. The benefit term Bdepends on the pathway flux J, while the cost term depends on the total pathway enzyme etot, and the fitness is given by the difference F=B−H=bJJ−γ etot with constant prefactors bJand γ. The cost His proportional to the enzyme level and may represent opportunity costs of the enzyme due to its ribosome and space demand. Note that fitness is measured here in arbitrary units, while the units of the coefficients bJand γmust such that cost and benefit functions have the same unit. For example, if the cost function denotes enzyme mass concentrations (with γ, in this case, being the enzyme molecular mass), then bJmust represent the benefit measured as an enzyme mass concentration per flux.
15 protein level (e) biomass rate (vBM) (A) Idle protein protein level (e) biomass rate (vBM) (B) Useful protein Figure 11.9: Protein/growth curves obtained from constraint-based metabolic models (schematic example) – A constraint-based metabolic model with fluxes and enzyme levels as variables defines a solution space, given by a highdimensional polytope. By projecting the polytope onto a plane spanned by a protein level of interest and biomass rate, we obtain a feasible set in protein/biomass space. If the original problem is a Linear Programming (LP) problem, this set is a convex polygon. Its silhouette curve defines the biomass/enzyme curve. In models with a fixed total protein amount, the biomass rate per protein amount can serve as a proxy for the cell growth rate (see Chapter 7 in [7]). The plots show, schematically, an idle protein (A) and a useful enzyme (B), where the optimum point is marked by an orange. 11.5.2. Predicting the effects of protein levels on growth in whole-cell models In the previous chapters, you saw cell models that–in principle–can predict protein levels that maximize cell growth. We can also use them to predict growth-rate effects of changes in protein levels. Let us have a look at some of these predictions! For simplicity, we consider linear resource allocation models, that is, models in which all rate laws are approximated by vl=elκlwith constant enzyme efficiencies1κl. In such models, a protein has two direct effects. First, it occupies a part the protein budget (described, for example, by a density constraint in the model) and thereby decreases the budget for other proteins and, in particular, for ribosomes. Second, if the protein is a metabolic enzyme, increasing its amount will allow the cell to increase metabolic fluxes and therefore precursor production. Both effects are indirect and may involve rearrangements of the proteome and of metabolic fluxes. In a cell model, these opposite effects require a best compromise, which can be determine in models by maximizing the growth rate. By screening the possible enzyme levels, we obtain curves relating enzyme levels to growth rates as in Figure 11.1. In practice, for constraint-based cell models the way to obtain such curves is always the same. A constraint-based model consists of a set of equalities and inequalities that define a set of solutions, each corresponding to a possible cell state with a certain cell growth rate and a certain expression level of the protein in question. By projecting this high-dimensional set onto a plane spanned by these two variables , we obtain a 2-dimensional set of possible protein/growth rate pairs. We may now assume that the cell, for each protein level, realizes the maximal possible growth rate. The resulting cell states will lie on the upper boundary of this set (see Fig. 11.9), a simple prediction of the curves in Figure 11.3. If our model is linear, the projected set is a convex polygon, with a bounding curve consisting of straight lines as shown in Figure 11.9. For an idle protein, the curve starts at a maximum at zero expression and decreases monotonically. For a useful protein, for example a metabolic enzyme used in the current conditions, the curve has a maximum at some positive expression level. In both cases, the decreasing part of the curve can be seen as a Pareto front describing a trade-off between our protein level and cell growth. In the case of a "useful protein", the curve has also an increasing branch on the left, describing a “win-win” situation in which higher enzyme levels allow for higher growth rates. 1In kinetic models (e.g. as described in chapter 6 in [7]), the κlvalues would be variable and dependent on metabolite levels. This would change the results, but here we assume that linear models capture the most important effects, including rearrangements of the proteome.
16 11.5.3. Minimal models trading metabolism against protein production As a simple example, we consider a variant of the first cell model in Chapter 8 in [7] (Section 8.3 in [7]). Metabolism and protein production are described by two overall reactions: EXT v1(e1) −−−−→ PRE v2(e2) −−−−→ MAC (11.8) The first reaction converts external metabolites (EXT) into precursors (PRE) and is catalyzed by an effective “metabolic enzyme” E1representing all metabolic enzymes in the cell. The second reaction converts the precursors into macromolecules (MAC) and is catalyzed by the “effective ribosome” E2, representing ribosomes, chaperones, tRNA, RNA polymerases and other machines, which we assume to act in fixed proportions. For simplicity, we denote both types of machines as "enzymes", even if ribosomes are actually RNA-protein complexes; and for a first simple model, we assume that the catalyzed reaction rates videpend linearly on the enzyme levels (v1=κ1e1and v2=κ2e2, with constant efficiencies κl). To relate the enzyme levels to the cell growth rate µ, we assume each of the reaction rates limits cell growth. This means that in optimal states the two rates must be equal (v1=v2) because otherwise, one of them would be higher than necessary and protein would be wasted. Hence, a steady state is already guaranteed without any further constraints. Assuming a fixed protein budget, we put a bound etot =ptot −q on the sum of enzyme levels, which leads to a trade-off between e1and e2. Here ptot is the total protein budget and qis the size of the "Q sector" of other, non-modeled proteins in the cell. Altogether, our model contains the variables v1, v2, e1, e2, µ and the parameters κl,etot, and q. Its feasible states must satisfy the constraints µ≤min(v1, v2) v1=κ1e1 v2=κ2e2 e1+e2+q≤ptot. (11.9) For each choice of parameter values (ptot, q, κ1, κ2), the constraints define a range of possible states (µ, v1, v2, e1, e2) (see Figure 11.10). The model (11.9) resembles the first model in Section 8.3 in [7] in Chapter 8 in [7] with two main differences. First, the cell growth rate is not directly given by the reaction rates, but only limited by them: (µ≤v1and µ≤v2). Second, we do not explicitly require a steady state; but since non-steady states with v16=v2 are suboptimal, the model predictions will be equivalent. As shown in Figure 11.10, the constraints (11.9) define a feasible polytope in the space spanned by e1, e2, and µ. The projection to the e1/µ plane yields a feasible triangle with a single growth-optimal point (see Figure 11.10(B)). What can we learn from this model? To simulate the expression of an idle protein, we assume that its abundance padds to the Q sector, thus decreasing the remaining budget for e1and e2(straight line in Figure 11.10(D)). In our model (11.9), the model variables µ,v1,v2,e1and e2scale proportionally: therefore, a decrease of the available enzyme budget e1+e2=ptot −qleads to a proportional decrease of the growth rate. With an idle protein adding to the Q sector, the growth rate changes as µ∼etot ptot −q=ptot −q−p ptot −q= 1 −p ptot −q.(11.10) or µ=µ01−p ptot −q.(11.11) This means that the protein/cost curve is a falling straight line that hits 0 where p=ptot −q, that is, where the idle protein takes up all the available protein budget, and no budget is left for the enzymes. In summary, our model predicts that useful proteins that contribute to the “metabolic enzyme” pool have a growth curve with an optimum
17 (A) min(ν1,ν2) ≥ μ e1 + e2 + q ≤ etot qe1 e1 e2 e2 (B) e1 e2 v1 e1 e2 v2 e1 e2 µ (C) e1 e2 µ e1 e2 µ e1 µ (D) p e2 µ p e2 µ p µ Figure 11.10: A minimal cell model can predict simple protein/growth curves – (A) Model with two reactions, catalyzed by two machines called "metabolic enzymes" and "effective ribosome". Each of the fluxes limits the cell growth rate, while the enzyme levels are limited by a fixed total protein budget. (B) According to Eq. (11.9), the reaction rates depend linearly on enzyme levels, and the growth rate µis limited by the minimum of the two rates, so the achievable growth rate µdepends on e1and e2. (C) Due to the protein bound, the solution space becomes a tetrahedron. Projecting it to the e1/µ plane yields a triangle (right) whose boundary is a piecewise linear protein/growth line. The curve consists of two straight lines with a maximum point in between. (D) For an idle protein, adding the protein level pto the Q sector and thereby decreasing the available enzyme budget leads to a single decreasing line with an optimum at zero expression, and zero cell growth where the idle protein takes up the entire available protein budget. If protein level and cell growth are seen as simultaneous optimization objectives, the falling part of the boundaries in (C) and (D) (marked in blue) forms a Pareto front. at a positive enzyme level. Idle proteins, in contrast, reduce the protein budget and therefore growth. The predictions assume a reallocation of protein resources between metabolism and translation, but they do not describe individual proteins with individual costs and benefits. To describe this, more detailed models are needed.
18 11.5.4. Flux-balance models with enzyme constraints Our simple model (11.8) predicts simple protein/growth curves, but it cannot describe metabolic enzymes individually. However, we can easily expand it by replacing the “effective metabolic reaction” by a metabolic network. We obtain more detailed protein sector models, as in Constrained Allocation FBA (CAFBA) [17], an FBA model with enzyme constraints and a ribosome sector (see chapter 5 in [7]). In such models, fluxes satisfy mass balance (N v = 0), catalytic constraints vi=κiei, and maybe heuristic flux bounds (e.g. positivity constraints vi≥0, after reorienting the reactions to impose realistic flux directions). We further assume a limited enzyme budget Plel≤etot and treat the biomass flux vBM (or a similar flux objective B=bJ·v) as the objective. Putting all this together, we can compute growth-optimal states. How will changes in enzyme parameters, such as catalytic rate kapp or molecule size, change the predicted growth rate? In FBA with enzyme constraints, each enzyme has two direct effects. On the one hand it catalyzes a flux v; on the other hand it occupies a part of the protein budget. In optimal states, the effect on the flux must be beneficial (i.e. the flux vimust have a positive marginal effect on the biomass rate, because otherwise the enzyme would not be expressed), and using the protein budget is costly (because of opportunity costs). We can now treat the enzyme level as a tunable parameter, optimize all other model variables, and compute the growth rate. By screening a range of enzyme levels, we obtain a parameter/growth curve, the upper line of the feasible set. At the optimal enzyme level, we obtain the optimal overall biomass/enzyme productivity–that is, the biomass production per enzyme–which can be translated into a predicted cell growth rate (see Chapter 7 in [7]). While FBA-like models are more detailed than our previous minimal model, the procedure for computing protein/- growth curves remains exactly the same. In the space of model variables, the model constraints define a polytope of valid states. Projecting this polytope onto a plane spanned by protein level and biomass rate leads to a feasible polygon. Its silhouette line describes the maximal possible biomass rate, satisfying all the model constraints, as a function of the protein level (see Figure 11.9). Qualitatively, the lines look like the curves in Fig. 11.5 or in the previous simple model. For idle proteins (or enzymes that do not contribute to a cost-efficient metabolic strategy in the given conditions), the line is strictly decreasing, with a maximum at zero expression. For useful proteins, the line has a maximum at a finite (non-zero) expression level. 11.5.5. Complex whole-cell models Resource Balance Analysis (RBA) models (see chapter 10 in [7]) are even more realistic than FBA. RBA models can describe macromolecular processes in great detail. Instead of assuming a given biomass composition, they optimize it together with the metabolic state, requiring that metabolism supplies all the precursors for macromolecules, which in turn are needed to produce other macromolecules or catalyze metabolic reactions. Hence, metabolism and macromolecule synthesis depend on each other and form a big feedback loop. Mathematically, possible cell states can be described in a space with the growth rate as one dimension and all other cell variables (fluxes and macromolecule concentrations) as the other dimensions. For each given growth rate, these other variables must satisfy linear constraints based on physical laws. Therefore, the space of valid cell states consists of high-dimensional polytopes, "stacked" along an extra, continuous growth rate dimension. To obtain protein/growth curves, we proceed as before: we project the set of feasible states (including the growth rate µas one of the variables) onto a plane spanned by µand the protein level in question. The main technical difference is that now we need to screen possible growth rate and determine, for each growth rate, a lower and upper bound for our protein level (where all other variables can be adjusted). However, just like before the silhouette line of the projected set yields the protein/growth rate curve (see Figure 11.11). Unlike in FBA, our model contains
19 Relative fitness Protein level [nmol/gCDW] 1 510 15 20 Ribosomes ATP synthase 0 20 40 60 0 0 Figure 11.11: Protein expression and growth rate effects predicted by a Resource Balance Analysis (RBA) model – Protein/growth curves for ATP synthase and ribosomes were computed by resource variability analysis in a genomescale RBA model of B. subtilis bacteria [18]. Cell growth, normalized to the maximal possible growth rate (and called "relative fitness") is shown on the x-axis and protein abundances are shown on the y-axis. The feasible sets were computed via resource variability analysis, similar to flux variability analysis in FBA (see chapter 5 in [7]) and indicate the achievable minimal and maximal values of ribosome and ATP synthase expression, assuming possible rearrangements of all cellular variables. The maximal ribosome amount decreases at higher growth rates, indicating a trade-off between ribosome amount and cell growth (similar for ATP synthase). Figure redrawn from [19]. nonlinear dilution rates of macromolecules and, possibly, growth rate-dependent parameters, vdil =µ cmacromol, so the silhouette lines may be curved. Similar to flux variability analysis in FBA (see chapter 5 in [7]), we may run a “resource variability analysis” to explore the allowed region and to compute the Pareto front (Figure 11.11). At each given growth rate (shown in a normalized form, as a relative fitness), a machine or a single protein can show a range of possible expression levels. At the maximum growth rate, this range will usually shrink to a point. To interpret the plot, it is good to remember that there is no fundamental difference between "variability analysis" and "multi-objective optimization". If a machine concentration xand the cellular growth rate are seen as objectives to be maximized, the boundary line yields a protein/growth curve as in Figure 11.1, and the part on the right can be seen as a Pareto front. In constraint-based models, the "objectives" (such as growth rate) and "model variables" (such as protein amounts) are not fundamentally different types of variables, but part of a large set of variables that mutually constrain each other. "Multi-objective optimization" is just a way to describe these effective constraints, and to present it as a trade-off between the two variables. So if we aim at predicting protein/growth curves, what do we gain by using RBA models? Like in FBA models, producing a useless protein will make the growth rate decrease, while for useful proteins, any deviation from its optimal, positive amount decreases the fitness. However, compared to FBA, the predictions account for more subtle effects and more complex adaptations of the cell, including changes in biomass composition. Unlike in FBA, varying protein levels will change the demands for biomass precursors, and therefore the required metabolic fluxes–which in turn lead to changing enzyme demands! For example, using proteins that contain trace elements such as metal ions will increase the demand for these ions. In contrast, when these ions are scarce, this has two different effects. First, it leads to a reallocation of protein towards pathways that scavenge these trace elements. But in addition, if ions are costly to obtain, the expression levels of the ion-containing proteins may be decreased. The logic may also hold for
20 nitrogen. Under low-nitrogen conditions in the environment, a simulated cell may both increase nitrogen import and decrease the use of proteins with a high nitrogen content. Hence, metabolic production and protein usage are tightly entangled via biochemical processes and cellular economics, and all this is reflected in the predicted protein/growth curves. 11.6. Protein and growth: insights from mechanistic models In this section, we stop taking for granted some of our previously made assumptions – most prominently, the notion that protein expression is optimal due to having been shaped by evolution – and consider the regulatory mechanisms actually making the cell behave according to them. Why should we now start questioning our notions? All models used by us throughout this chapter to understand the effects of protein expression rely on constraints, which define the set of possible cell states and thus the set of cell behaviors realistically achievable in real life. However, not all constraints are of the same nature. Some stem from the first principles of molecular biology, such as the non-negativity of concentrations or mass conservation constraints in FBA and RBA. Others, rather than addressing the question of what rules the cell must obey, focus on how the cell tends to behave. Such constraints are called “phenomenological” since they capture commonly observed phenomena without a fundamental explanation. An instance of this is the assumption of linear correlation between the cell’s growth rate and its overall ribosome abundance (protein synthesis budget) often made in minimal constraint-based models [20], which reflects the trends observed in experimental data. In this vein, optimality may be considered the ultimate phenomenological constraint, with frameworks like FBA and RBA relying on it to predict the cell’s state simply because growth rate maximization is often observed in nature. However, in some cases, these phenomenological constraints may be invalid, as no inherent properties of the cell prevent their violation. Looking at the effects of protein expression in particular, at high protein production rates, the relationship between the ‘useless’ protein mass fraction and the growth rate deviates from the straight line [21] predicted by minimal constraint-based models [20] (e.g. the upper boundary in Figure 11.12D) . Moreover, the regulation of certain metabolic proteins’ expression in E. coli has been experimentally determined to be suboptimal [22]. Suboptimal scenarios may also arise – and be particularly common – when we engineer cells with synthetic proteins, as the cell, by definition, has not had time to evolve and optimize their expression. Since phenomenological constraints can sometimes be misleading, instead of enforcing commonly observed cell behaviors (which often means best-possible behaviors), it can be useful to ask: “what are the actual cellular regulation mechanisms enabling these phenomena?” To tackle this question, we can use mechanistic ordinary differential equation (ODE) models, which capture the dynamics of different cellular variables. These models still use ‘first-principle’ constraints (like the laws of enzyme kinetics), but bake them into the form of the differential equations rather than enforce them explicitly. The equations also incorporate the terms for particular molecular regulation mechanisms believed to make the cell closely (but possibly not perfectly) adhere to the phenomenological constraints. Due to the complexity of living systems, such ODE models are usually coarse-grained, with each variable describing the average dynamics of multiple biomolecules with similar functions and properties. The degree of this coarse-graining can be varied based on the desired level of detail and accuracy, as well as the scenario considered. In many cases, mechanistic models reproduce constraint-based modeling predictions, either by simulation or analytically. The latter involves solving a model for the steady state (i.e. setting all ODEs to zero) and rearranging the terms of the equations. For example, models based on the Flux-Parity Regulation theory [21,23], which capture the regulation of ribosomes by the ppGpp signaling molecule, whose abundance is proportional to the ratio between charged and uncharged tRNA concentrations, algebraically yield the same linear relation between useless protein expression and its cost to the cell (see Box 11.D and Figure 11.12).
21 Box 11.D Deriving the cost of useless protein expression To derive a mathematical expression for the cost of expressing a protein without any beneficial influence on cell growth, we start by defining a simple mechanistic E. coli cell model from first principles, which is given by Ordinary Differential Equations (11.12) and proteome allocation relations from Equations (11.13). The model captures the total protein biomass of a growing cell population (M) and the uncharged and aminoacylated tRNA concentrations per cell (Tuand Tc, respectively). The fractions of the cell’s proteome taken up by the ribosomal proteins, carbon metabolism proteins, housekeeping proteins (native proteins not in the two previous sectors) and the useless protein are respectively denoted as φr,φa,φqand φx. The ODEs below reflect the following modeling assumptions, which have been informed by experimental studies [23,21]: ◦The total protein mass density per unit cell volume, M, is constant, so the rate of the tRNA species’ dilution due to cell growth (i.e. volume expansion) equals the rate of biomass increase λ. This also means that the protein mass fractions are proportional to protein concentrations per unit cell volume. ◦tRNA aminoacylation, consuming uncharged and producing charged tRNAs, is catalyzed by metabolic proteins (hence its dependence on φa). Similarly, translation (and thus protein biomass synthesis) is catalyzed by ribosomes and depends on the abundance of aminoacyl-tRNAs. Each reaction’s rate exhibits a MichaelisMenten dependence on the concentration of its precursor. ◦The signaling molecule ppGpp, whose level is given by the ratio of charged and uncharged tRNA abundances, controls proteome allocation between ribosomal and metabolic gene. This is captured by a Hill function. ◦The rate of RNA transcription, including that for tRNAs, is proportional to the cell’s growth rate. dM dt =γTc Ktrans +Tc φrM=λM dTc dt =νTu Kaa +Tu φa−γTc Ktrans +Tc φr−λTc dTu dt =ψ(Tc/Tu) KppGpp +Tc/Tu ·λ−νTu Kaa +Tu φa+γTc Ktrans +Tc φr−λTu(11.12) where φq=const, φx=const, φr= (1 −φq−φx)·Tc/Tu KppGpp +Tc/Tu , φa= (1 −φq−φx)·KppGpp KppGpp +Tc/Tu (11.13) For a cell in the steady state, tRNA concentrations are in equilibrium. Taking dTc dt = 0 and dTu dt = 0 and substituting the definitions from Equation (11.13), we can observe that all φx-dependent terms cancel out. Hence, charged and uncharged tRNA abundances in the steady state are predicted to be unaffected by the useless protein’s expression. Then, the steady-state cost of having the useless protein occupy the fraction φ0 xof the cell’s proteome can be found according to Equation (11.14), where we use the fact that the φx-independent tRNA abundances Tuand Tcare the identical in the numerator and the denominator of the fraction. η(φ0 x) = λ(φx= 0) −λ(φx=φ0 x) λ(φx= 0) = 1 − γTc Ktrans+Tc·(1 −φq−φ0 x)·Tc/Tu KppGpp+Tc/Tu γTc Ktrans+Tc·(1 −φq−0) ·Tc/Tu KppGpp+Tc/Tu ⇔ ⇔η(φ0 x) = 1 −1−φq−φ0 x 1−φq−0=φ0 x 1−φq (11.14) This linear increase in costs as the useless protein’s abundance in the cell increases is likewise predicted by the constraint-based model in Section 11.5.3. In the present case, however, we made no assumption of optimality but rather considered the experimentally determined growth regulation mechanisms via ppGpp signaling. This relation is also concordant with experimental data (see Figure 11.12) for sufficiently low φxvalues. Greater useless protein abundances, which yield a significant divergence from this law, are theorized to activate the cellular stress response mechanisms neglected by the simple cell model in Equations (11.12)–(11.13) [23].
22 0 0.1 0.2 0.3 0.4 0 20 40 60 80 100 Excess protein mass frac. in the cell Cost, η(relative growth rate reduction) Exp. data Model Figure 11.12: Experimentally measured synthetic protein mass fractions in the cell plotted against cell growth rates relative to cells expressing no useless protein (data taken from [21]). The red line represents the linear dependency observed in Figure 11.10 and derived from a mechanistic model in Box 11.D. At higher mass fractions, this linear relation is in practice violated, with costs growing superlinearly, due to the phenomena neglected by the mechanistic model. In more detailed models, besides drawing from the ribosomal budget, protein expression is assumed to have positive and negative effects on other aspects of cellular functioning, such as its ATP and amino acid consumption and synthesis, all captured by different parameters. Varying these parameters individually and in conjunction yields different curve shapes, which can be compared to experimentally observed trends. This may help to deduce how exactly a given protein favors and hinders growth by looking at this protein’s cost-benefit curves [24]. What if predictions from an ODE model disagree with reality, in the same way that constraint-based models’ phenomenological constraints are sometimes violated? As discussed in Chapter 12 in [7], such observations provide valuable information, revealing that there is a regulatory mechanism or a biomolecular reaction currently overlooked by us. Namely, for the cost of expressing the LacZ or ∆tufB proteins, experimental measurements diverged significantly from the predictions obtained by simulating a mechanistic cell model from [24]. This discrepancy, however, was removed by redefining the protein degradation rate, changing it from a simple linear relationship to a Hill formula. The non-linearity was explained by noting that bacterial proteases primarily degrade misfolded proteins; hence, the need to model the dynamics of protein folding and its dependence on the overall protein expression levels. 11.7. Concluding remarks Any process in the cell can be viewed through the lens of its effect on the cell’s fitness under given environmental conditions. By “fitness” we usually mean the efficiency with which the cell can increase its biomass – that is, the rate of cell growth. Furthermore, growth can easily be related to fluxes of matter in metabolic models we often use for predictions, . Considering growth rates also facilitates the economic analogy between a company’s drive to increase its market value and the cell’s incentive to build up biomass. This is the view we have adopted in this chapter; however, considerations would be similar for other proxy variables capturing a cell’s fitness. A process of prime importance for cell growth is protein expression, since proteins enable the majority of metabolic reactions and comprise the greatest share of the cell’s biomass, while their synthesis consumes most of the cellular energy (i.e. ATP molecules) and other resources. In this chapter, therefore, we discussed the question: how to determine the fitness effects of expressing a given protein? This consideration is crucial for predicting the cell’s metabolic regulation and
23 (A) protein fitness effect F=fitness benefit B−fitness cost H (B) Steady-state protein expression Fitness effect Expression usually assumed optimal: .marginal benefit % and -marginal cost & cancel out Likely causes of non-optimality: protein is non-native cell is a non-model organism protein function depends on other unconsidered proteins Figure 11.13: The fitness effects of expressing a protein are (A) calculated as the difference between the protein’s fitness and cost and (B) are usually assumed to be driven to an optimum by evolution. understanding how it has evolved, as well as for engineering the expression of heterologous proteins in cells to be optimal. As we saw in this chapter, it is useful to separately consider the protein’s positive and negative contributions to the growth rate – that is, its benefit and cost to the cell – and represent the protein’s overall fitness effects as the difference of these terms (see Figure 11.13). This (artificial, rather than inherently physical) distinction allows one to draw parallels with economics, in which any activity’s ultimate profitability for a company is calculated by subtracting the costs of undertaking it from its contribution to the revenue. The possibility of defining the overall effect as a linear difference is a simplifying assumption we make; however, it is justifiable by taking a linear approximation when operating with small variations in gene expression and is a natural property of certain metabolic models, such as those in which the cell’s overall protein abundance is fixed. In other cases, a linear relationship can be obtained by redefining the fitness function, e.g. taking a logarithm of the ratio between the metabolic flux facilitated by the enzyme (its benefit) and its total amount (proportional to its cost) yields a difference between log-terms. Another convenient assumption is that the protein’s abundance is the sole argument of the cost and benefit functions, which makes the effects of disparate proteins comparable to each other. A protein’s costs and benefit may be understood as purely empirical values relating the observed protein levels, varied by the experimentalist, to the cell’s growth rates as measured using techniques discussed in the appendix. This approach is illustrated by the Lac operon case study, where first the cost is established in an experiment in which the protein of interest is rendered useless to the cell (by not having the enzyme’s substrate present). Afterward, the benefit is determined by observing changes in cell growth for fixed protein expression levels with known costs. However, costs and benefits can also be treated more mechanistically by considering how they arise. Namely, benefits stem from proteins carrying out their biological functions. Meanwhile, expressing a given protein has direct “opportunity” costs – that is, different resources required for its synthesis and the space it occupies in the cell’s proteome are unavailable for other proteins that are potentially necessary for growth. Moreover, changes in a protein’s expression
24 bring about global changes in the highly interconnected cellular metabolic network. Most prominently, increasing a protein’s production raises demand for the cell’s expression machinery, and increasing ribosome expression to meet this demand will itself consume more resources. Such effects may be called indirect costs, or “baggage”. Defining mathematical models with different cost and benefit relations built in – potentially dependent on a given protein’s molecular properties and enzyme efficiency – and comparing their predictions with experimental data can elucidate the fitness effects of expressing the protein. Appreciating the costs and benefits of protein expression can help us understand how the cell manages its proteins and the evolutionary pressures that have shaped this behavior. If we assume that evolution drives a protein’s expression towards enabling optimal fitness, its level can be predicted by retrieving a “sweet spot” between its cost and benefit. At this point, marginal (i.e. per one additional unit of protein produced) costs and benefits must cancel each other out, so that neither lowering nor increasing expression can improve fitness. Since costs per unit protein always increase, the existence of such an optimal level is conditional on the marginal benefits being positive. At the same time, to have an optimal protein level which is finite, protein expression must be constrained or must bring about diminishing returns as it increases further and further. Nonetheless, care must be taken when applying the cost-benefit paradigm described in this chapter. Namely, optimization of gene expression for maximum fitness is merely an assumption backed by the fact that fastest-reproducing cells are usually favored by evolution. Hence, some proteins’ expression may be near-optimal in some conditions but still be found away from the supposed “sweet spot” in others. To explain this, one can consider models which explicitly incorporate the mechanisms of gene regulation, as well as protein expression and processing, noting which modeling assumptions produce good agreement with experimental data. Non-optimality may be particularly likely when the cell is engineered with synthetic genes encoding new proteins whose expression, unlike that of the cell’s natural genes, has not yet been optimized by evolution. However, if proteins fitness effects are studied and incorporated into models, their predictions can help to achieve such synthetic protein expression that cell population growth or production of desired compounds are maximized [23,25]. An important caveat in all the above considerations is that costs and benefits are highly context-dependent. For instance, while almost any increase in a protein’s expression in E. coli (which we focus on here) involves an observable cost, the same may not always be true for other organisms, such as S. cerevisiae [5,6] or B. subtilis [18]. This is likely to stem from the fact that, depending on the conditions, some cells do not optimize their protein expression for maximum instantaneous growth, but rather maintain spare protein synthesis capacity to ensure adaptability to changing circumstances [26,27,19]. Within the cell, context-dependence may arise if the protein of interest operates in a complex with other proteins or is part of a metabolic pathway. Namely, for the latter, the expressed protein’s fitness effects depend not only on substrate concentrations, but also on the levels of other proteins in the same pathway, which may decide whether upregulating the protein has positive marginal benefits by boosting the biomass flux or whether the entire pathway is too wasteful and the protein always has negative fitness effects. Recommended readings ◦Empirical cost and benefit functions and evolution towards predicted optimal expression. Erez Dekel and Uri Alon. Optimality and evolutionary tuning of the expression level of a protein. Nature, 436(7050):588592, 2005. doi: 10.1038/nature03842. ◦A comprehensive description of metabolic control and optimality on the level of entire cells. Frank J. Bruggeman, Maaike Remeijer, Maarten Droste, Luis Salinas, Meike Wortel, Robert Planqué, Herbert M. Sauro, Bas Teusink, and Hans V. Westerhoff. Whole-cell metabolic control analysis. BioSystems, 234:105067, 2023. doi: 10.1016/j.biosystems.2023.105067
31 [14] Jens G. Reich. Zur Ökonomie im Proteinhaushalt der lebenden Zelle. Biomed. Biochim. Acta, 42(7/8):839–848, 1983. [15] I. Shachrai, A. Zaslaver, U. Alon, and E. Dekel. Cost of unneeded proteins in E. coli is reduced after several generations in exponential growth. Molecular Cell, 38:1–10, 2010. doi: 10.1016/j.molcel.2010.04.015. [16] Hisao Moriya, Yuki Shimizu-Yoshida, and Hiroaki Kitano. In vivo robustness analysis of cell division cycle genes in Saccharomyces cerevisiae. PLoS genetics, 2(7):e111, 2006. doi: 10.1371/journal.pgen.0020111. [17] Matteo Mori, Terence Hwa, Olivier C. Martin, Andrea De Martino, and Enzo Marinari. Constrained allocation flux balance analysis. PLoS computational biology, 12(6):e1004913, 2016. doi: 10.1371/journal.pcbi.1004913. [18] Anne Goelzer, Jan Muntel, Victor Chubukov, Matthieu Jules, Eric Prestel, Rolf Nölker, Mahendra Mariadassou, Stéphane Aymerich, Michael Hecker, Philippe Noirot, et al. Quantitative prediction of genome-wide resource allocation in bacteria. Metabolic engineering, 32:232–243, 2015. doi: 10.1016/j.ymben.2015.10.003. [19] Oliver Bodeit, Inès Ben Samir, Jonathan R. Karr, Anne Goelzer, and Wolfram Liebermeister. Rbatools: a programming interface for resource balance analysis models. Bioinformatics Advances, (vbad056), 2023. [20] Matthew Scott, Carl W Gunderson, Eduard M Mateescu, Zhongge Zhang, and Terence Hwa. Interdependence of cell growth and gene expression: origins and consequences. Science, 330(6007):1099–1102, 2010. doi: 10.1126/science.1192588. [21] Griffin Chure and Jonas Cremer. An optimal regulation of fluxes dictates microbial growth in and out of steady state. eLife, 2023. doi: 10.7554/eLife.84878. [22] Benjamin D. Towbin, Yael Korem, Anat Bren, Shany Doron, Rotem Sorek, and Uri Alon. Optimality and suboptimality in a bacterial growth law. Nature Communications, 8:article number 14123, 2017. doi: 10.1038/ ncomms14123. [23] Kirill Sechkar, Harrison Steel, Giansimone Perrino, and Guy-Bart Stan. A coarse-grained bacterial cell model for resource-aware analysis and design of synthetic gene circuits. Nat. Commun., 15(1981):1–17, 2024. doi: 10.1038/s41467-024-46410-9. [24] Chen Liao, Andrew E. Blanchard, and Ting Lu. An integrative circuit–host modelling framework for predicting synthetic gene network behaviours. Nat. Microbiol., 2:1658–1666, 2017. doi: 10.1038/s41564-017-0022-5. [25] François Bertaux, Jakob Ruess, and Grégory Batt. External control of microbial populations for bioproduction: A modeling and optimization viewpoint. Current Opinion in Systems Biology, 28:100394, 2021. doi: 10.1016/ j.coisb.2021.100394. [26] Moshe Kafri, Eyal Metzl-Raz, Ghil Jona, and Naama Barkai. The cost of protein production. Cell Rep., 14(1): 22–31, 2016. doi: 10.1016/j.celrep.2015.12.015. [27] Manlu Zhu, Qian Wang, Haoyan Mu, Fei Han, Yanling Wang, and Xiongfeng Dai. A fitness trade-off between growth and survival governed by Spo0A-mediated proteome allocation constraints in Bacillus subtilis. Sci. Adv., 9(39), 2023. doi: 10.1126/sciadv.adg9733. [28] Nguyen Quoc Khanh Le, Edward Kien Yee Yapp, N. Nagasundaram, and Hui-Yuan Yeh. Classifying promoters by interpreting the hidden information of DNA sequences via deep learning and combination of continuous FastText N-grams. Frontiers in Bioengineering and Biotechnology, 7, 2019. doi: 10.3389/fbioe.2019.00305. [29] Travis L. LaFleur, Ayaan Hossain, and Howard M. Salis. Automated model-predictive design of synthetic promoters to control transcriptional profiles in bacteria. Nature Communications, 13(1), 2022. doi: 10.1038/ s41467-022-32829-5.
32 [30] Heidi Redden and Hal S. Alper. The development and characterization of synthetic minimal yeast promoters. Nature Communications, 6(1), 2015. doi: 10.1038/ncomms8810. [31] Michael E. Lee, Anil Aswani, Audrey S. Han, Claire J. Tomlin, and John E. Dueber. Expression-level optimization of a multi-enzyme pathway in the absence of a high-throughput assay. Nucleic Acids Research, 41(22): 1066810678, 2013. doi: 10.1093/nar/gkt809. [32] Tim Weenink, Jelle van der Hilst, Robert M McKiernan, and Tom Ellis. Design of RNA hairpin modules that predictably tune translation in yeast. Synthetic Biology, 3(1), 2018. doi: 10.1093/synbio/ysy019. [33] Howard M Salis, Ethan A Mirsky, and Christopher A Voigt. Automated design of synthetic ribosome binding sites to control protein expression. Nature Biotechnology, 27(10):946950, 2009. doi: 10.1038/nbt.1568. [34] Lior Zelcbuch, Niv Antonovsky, Arren Bar-Even, Ayelet Levin-Karp, Uri Barenholz, Michal Dayagi, Wolfram Liebermeister, Avi Flamholz, Elad Noor, Shira Amram, Alexander Brandis, Tasneem Bareia, Ido Yofe, Halim Jubran, and Ron Milo. Spanning high-dimensional expression space using ribosome-binding site combinatorics. Nucleic Acids Research, 41(9):e98e98, 2013. doi: 10.1093/nar/gkt151. [35] Iman Farasat, Manish Kushwaha, Jason Collens, Michael Easterbrook, Matthew Guido, and Howard M Salis. Efficient search, mapping, and optimization of multiprotein genetic systems in diverse bacteria. Molecular Systems Biology, 10(6), 2014. doi: 10.15252/msb.20134955. [36] Elisa Marquez-Zavala and Jose Utrilla. Engineering resource allocation in artificially minimized cells: Is genome reduction the best strategy? Microbial Biotechnology, 16(5):990999, 2023. doi: 10.1111/1751-7915.14233. [37] Victoria Munro, Van Kelly, Christoph B. Messner, and Georg Kustatscher. Cellular control of protein levels: A systems biology perspective. PROTEOMICS, 24(1213), 2023. doi: 10.1002/pmic.202200220. [38] Fredrik Edfors, Frida Danielsson, Björn M Hallström, Lukas Käll, Emma Lundberg, Fredrik Pontén, Björn Forsström, and Mathias Uhlén. Gene-specific correlation of RNA and protein levels in human cells and tissues. Molecular Systems Biology, 12(10), 2016. doi: 10.15252/msb.20167144. [39] Yun-Chi Tang and Angelika Amon. Gene copy-number alterations: A cost-benefit analysis. Cell, 152(3):394405, 2013. doi: 10.1016/j.cell.2012.11.043. [40] Francesca Ceroni, Rhys Algar, Guy-Bart Stan, and Tom Ellis. Quantifying cellular capacity identifies gene expression designs with reduced burden. Nature Methods, 12(5):415418, 2015. doi: 10.1038/nmeth.3339. [41] Idan Frumkin, Dvir Schirman, Aviv Rotman, Fangfei Li, Liron Zahavi, Ernest Mordret, Omer Asraf, Song Wu, Sasha F. Levy, and Yitzhak Pilpel. Gene architectures that minimize cost of gene expression. Molecular Cell, 65 (1):142153, 2017. doi: 10.1016/j.molcel.2016.11.007. [42] Guillaume Cambray, Joao C Guimaraes, and Adam Paul Arkin. Evaluation of 244, 000 synthetic sequences reveals design principles to optimize translation in Escherichia coli. Nature Biotechnology, 36(10):10051015, 2018. doi: 10.1038/nbt.4238.