The Millennium Problems for Biology
Recorded: Sept. 20, 2026, 1:09 p.m.
| Original | Summarized |
The Millennium Problems for BiologyEdison Scientific · FutureHouseThe Millennium Problems for Biology01Origins of life+Demonstrate the emergence of life from chemical precursors in a laboratory setting.Specifically, demonstrate the unassisted emergence of self replicating RNA- and protein-based cells from a plausible primordial soup with a plausible energy source. A “cell” may be any compartment with a defined boundary. To be considered successful, the following conditions must be met. Firstly, it must be shown that the emergent cells can increase their abundance by at least a factor of 10⁶ (roughly 20 generations), when provided with sufficient primordial soup and energy. Secondly, it must be plausible that division could continue indefinitely given sufficient energy and primordial soup. For example, solutions that involve the cells monotonically decreasing in size over successive divisions would not be accepted. Finally, the cells must have a clear way of encoding heritable genetic information, i.e., the molecular composition of the cells must be causally determined at least in part by information stored within the cell. Solutions that involve storing the information in the form of nucleic acids, polypeptides, or similar polymers are strongly preferred. Solutions in which the existence of heritable genetic information is ambiguous or controversial will be rejected by default.02Cryopreservation+Demonstrate the ability to cryopreserve and recover live wild-type mice with high viability.Specifically, demonstrate the reversible cryopreservation of live, intact, wild-type adult mice in a whole-body frozen or vitrified state. The mice must remain frozen or vitrified for at least 24 hours, must be recovered with >99% viability, and must not suffer any permanent organ damage or bodily harm. Somatic genetic engineering is discouraged but permitted. All experiments must be conducted with ethics approval.03The reverse translatase+Create an enzyme that can “reverse translate” an arbitrary peptide sequence into RNA or DNA.Specifically, create a purified protein catalyst or fixed protein complex that processively reads an untagged polypeptide and synthesizes a covalent nucleic acid strand encoding its residue sequence under a preregistered codon convention, without a nucleic-acid template, preattached sequence barcode, residue-specific operator cycle, or database lookup. The resulting nucleic acid strand must be compatible with ordinary polymerases, ligases, and other similar enzymes, i.e., if nucleic acids other than RNA or DNA are used, they must be compatible with downstream amplification or sequencing reactions. For the challenge to be considered complete, at least 100 random peptide sequences of at least 50 amino acids each must be preregistered, synthesized, and pooled. It must then be shown that the sequences of these peptides can be inferred, without reference to a dictionary, by reverse translation and sequencing with at least 90% sequence accuracy. Moreover, the average read length must be at least 25 residues, and the average read quality score should be at least Q10.04Improve Rubisco+Produce a Rubisco enzyme with specificity and enzymatic turnover beyond the naturally occurring pareto frontier.Specifically, produce an enzyme that catalyzes the carboxylation of ribulose-1,5-bisphosphate with a specificity for carbon dioxide over oxygen (Sc/o) at least as high as that of Galdieria Partita Rubisco, and with an enzymatic turnover (kcat) at least as high as that of maize Rubisco. To be considered successful, the specificity and enzymatic turnovers of the candidate enzyme must be measured in paired enzyme assays using G. Partita Rubisco and maize Rubisco as controls, respectively. The candidate enzyme may be designed de novo, discovered in nature, or engineered or evolved from naturally occurring starting points.05The quadruplet cell+Produce a living cell that uses a four-base codon code.Specifically, produce a living and replicating cell in which every protein-coding sequence, including the translation machinery itself, is encoded as uninterrupted nonoverlapping quadruplet codons, without detectable triplet decoding. The encoding scheme must be a bona fide quadruplet encoding, i.e., in the quadruplet encoding, the probability that a mutation is non-synonymous must be similar regardless of the index of the mutation in the codon. For example, quadruplet encodings in which the first three codon positions are always or almost always sufficient to specify the encoded amino acid will not be accepted.(Contributed by Erika Alden DeBenedictis)06Somatic limb regeneration+Demonstrate the ability to regenerate lost limbs in adult wild-type mice.Specifically, demonstrate, in an adult wild type mouse, the reproducible ability to regrow limbs following amputation. Following regeneration, the mouse must perform indistinguishably from controls in a standard battery of motor function tests, must demonstrate indistinguishable sensory perception in the regrown limb, and blinded observers must not be capable of distinguishing which limb was regrown based on non-invasive observational data. All experiments must be conducted with ethics approval.07Bacterial production of gene therapies+Demonstrate the ability to produce gene therapies in a bacterial host.Specifically, produce infectious replication-incompetent AAV and lentivirus in bacteria. The particles must contain a pre-specified viral genome; the ratio of physical capsids to viral genomes and the ratio of infectious units to viral genomes must be similar to the ratios obtained when purifying viruses from mammalian cell culture; and the viral genomes must be nuclease-resistant. It is anticipated that producing lentivirus in bacteria may be much more challenging than producing AAV, and thus demonstrating the ability to produce AAV on its own will be considered a partial success.08Programmable proteases+Demonstrate the ability to produce enzymes on demand that will specifically and efficiently cut a specific protein sequence.Specifically, given a blinded, accessible site in an endogenous folded protein, demonstrate the ability to prospectively design a protease that cleaves that site efficiently in living cells. The resulting enzyme must have catalytic efficiency and proteome-wide off-target cleavage similar to or greater than other widely-used site-specific proteases. The challenge will be considered complete when the design can be demonstrated against 20 preregistered sites with a success rate greater than 80%. Once the target sites are preregistered, the designs of the resulting proteins must be produced within 24 hours, and no wet lab work is allowed prior to evaluation except for the purpose of producing the designed proteins for assay. (Hence, for example, screening and target-specific evolution are not permitted once the target sites are provided.)Note that a weaker form of this challenge involves demonstrating the ability to produce enzymes that specifically and efficiently cleave specific preregistered peptide sequences, when those sequences are provided in solution, along with off-target sequences. Demonstration of that ability will be considered a partial success.09Cell-penetrating protein binders+Demonstrate the ability to produce protein binders against intracellular targets.Specifically, demonstrate the ability to design zero-shot protein binders that, without further evolution or optimization, will reliably engage preregistered intracellular protein targets in living cells when administered extracellularly to those cells at pharmacologically supported concentrations. The cell entry mechanism must be plausible in a therapeutic context, i.e., transfection, intrabody expression, electroporation, membrane disruption, or similar methods are not permitted. The challenge will be considered complete when the design can be demonstrated against 20 preregistered targets with an 80% success rate. Once the targets are preregistered, the designs of the resulting proteins must be produced within 24 hours, and no wet lab work is allowed prior to evaluation except for the purpose of producing the designed proteins for assay. (Hence, for example, screening and target-specific evolution are not permitted once the targets are provided.)The original intention of this problem was specifically to design antibodies against intracellular targets. However, it is anticipated that modifications to the antibody scaffold will be required in order for the problem to be solvable. Since we cannot put an upper bound on the magnitude of the modifications required, we have broadened the problem to encompass any protein binders. However, solutions that involve binders resembling humanized monoclonal antibodies will be greatly preferred. The problem would likely be even more impactful if solved in general for small molecule binders, rather than protein binders or antibodies. However, with small molecule binders, synthesis is a major bottleneck that would limit validation, and thus we have chosen to restrict the scope to protein binders.10Protein amplification chain reaction+Demonstrate exponential amplification of arbitrary peptide substrates.Specifically, demonstrate input-protein-dependent synthesis of new, full-length, sequence-faithful covalent polypeptide copies from amino-acid monomers without a nucleic-acid template or preformed cognate scaffold, in a single pot reaction. For the challenge to be considered complete, at least 100 random peptide sequences of at least 50 amino acids each must be preregistered, synthesized, and pooled. It must then be shown that the abundance of these peptides in solution can be amplified at least 1000x with at least 90% sequence accuracy on a per-residue basis. Reasonable modifications may be added to the peptide sequences to facilitate post-amplification analysis if necessary, provided they are not active in the amplification. Methods that rely on explicit sequencing of the peptide are not permitted. Methods that rely on reverse translation to generate a nucleic acid intermediate are not permitted, because they are duplicative with a separate Millennium Problem.115′ polymerases+Produce a full set of polymerases that act in the 5′ direction.Specifically, produce a complete set of 3′>5′ polymerases comparable to commonly used 5′>3′ polymerases, including a 3′>5′ DNA polymerase, a 3′>5′ RNA polymerase, a 3′>5′ reverse transcriptase, and a 3′>5′ RDRP. The proteins should have processivity and error characteristics that are similar to or superior to those of Taq, T7 RNA pol, M-MLV RT, and Phi 6 RDRP respectively. These proteins may be designed de novo, discovered in nature, or engineered or evolved from naturally occurring starting points.12New nitrogenases+Create a new nitrogenase that does not bear sequence or structural homology to the natural family.Specifically, the protein must convert N₂ to ammonia at rates that are at least of a similar order of magnitude to the rates of naturally occurring proteins, and must fall well below the sequence- and structure-similarity thresholds relative to all known nitrogenase and nitrogenase-like proteins. The protein may be designed de novo, discovered in nature, or engineered or evolved from naturally occurring starting points.Sam Rodriques · Michaela Hinks |
The Millennium Problems for Biology outline twelve complex scientific challenges demanding innovative solutions across diverse fields. The first challenge addresses the origins of life, requiring the demonstration of how life can emerge from chemical precursors in a laboratory setting by creating self-replicating RNA- and protein-based cells from a primordial soup and an energy source. Success requires that these emergent cells can increase their abundance by at least a factor of ten to the power of six over approximately twenty generations, can divide indefinitely with sufficient resources, and possess a mechanism for encoding heritable genetic information, preferably in nucleic acids or similar polymers. A second challenge concerns cryopreservation, focusing on demonstrating the reversible cryopreservation of live, intact, wild-type adult mice in a whole-body frozen or vitrified state, ensuring they maintain over ninety-nine percent viability without suffering permanent organ damage, while permitting somatic genetic engineering under ethical supervision. The third problem involves the creation of a reverse translatase enzyme capable of translating an arbitrary peptide sequence into an RNA or DNA strand. This enzyme must processively read an untagged polypeptide and synthesize a covalent nucleic acid strand encoding the sequence adhering to pre-registered codon conventions without relying on a nucleic acid template or sequence barcode. Completing this endeavor involves preregistering at least one hundred random peptide sequences of at least fifty amino acids each, and then demonstrating that these sequences can be inferred with at least ninety percent sequence accuracy through reverse translation and sequencing. Improving Rubisco is the fourth challenge, which seeks to produce an enzyme with specificity and enzymatic turnover that exceeds the natural pareto frontier. The goal is to engineer an enzyme that catalyzes the carboxylation of ribulose-1,5-bisphosphate with a specificity for carbon dioxide over oxygen that is at least as high as that of Galdieria Partita Rubisco, and an enzymatic turnover rate comparable to that of maize Rubisco, verified via paired enzyme assays. The fifth challenge involves developing a living cell that utilizes a four-base codon code for encoding all protein-coding sequences, including the translation machinery, without the need for detectable triplet decoding. This encoding scheme must be a valid quadruplet encoding, ensuring that mutation probability for non-synonymous changes is consistent regardless of the mutation's position within the codon. The sixth challenge addresses somatic limb regeneration in adult wild-type mice, requiring the reproducible ability to regrow lost limbs following amputation. This regeneration must result in a mouse that performs motor functions and sensory perceptions indistinguishable from controls, and observers must be unable to distinguish the regrown limb based on non-invasive data. The seventh problem focuses on bacterial production of gene therapies, necessitating the ability to produce replication-incompetent adeno-associated virus (AAV) and lentivirus within bacteria. The resulting particles must contain a pre-specified viral genome, exhibiting ratios of physical capsids to viral genomes and infectious units to viral genomes similar to those obtained from mammalian cell culture purification, and the viral genomes must be resistant to nuclease degradation. The eighth challenge calls for programmable proteases, aiming to create enzymes that can be produced on demand to specifically and efficiently cleave a designated protein sequence within living cells. The designed enzyme must possess catalytic efficiency and off-target cleavage properties similar to or better than established site-specific proteases, achieving success against twenty preregistered sites with an eighty percent success rate. The ninth problem targets cell-penetrating protein binders, requiring the design of zero-shot protein binders that reliably engage pre-registered intracellular protein targets in living cells when administered extracellularly. The mechanism for cell entry must be plausible in a therapeutic context, and the design must demonstrate success against twenty pre-registered targets with an eighty percent success rate. The tenth challenge focuses on a protein amplification chain reaction, demanding the demonstration of exponential amplification of arbitrary peptide substrates. This process must involve the input of protein-dependent synthesis of new, full-length, sequence-faithful covalent polypeptide copies from amino-acid monomers without requiring a nucleic acid template or reverse translation. The amplification must achieve at least one thousandfold increase in abundance with at least ninety percent sequence accuracy per residue, without relying on explicit sequencing methods. The eleventh problem requires the production of a complete set of polymerases that operate in the five prime direction, including a three prime to five direction DNA polymerase, RNA polymerase, reverse transcriptase, and reverse DNA polymerase. These proteins should exhibit processivity and error characteristics comparable to or superior to naturally occurring counterparts such as Taq, T7 RNA polymerase, M-MLV reverse transcriptase, and Phi 6 reverse DNA polymerase, respectively. Finally, the twelfth challenge seeks the creation of a novel nitrogenase that lacks any sequence or structural homology to the natural family of nitrogenase and nitrogenase-like proteins, yet retains the ability to convert dinitrogen into ammonia at rates comparable to naturally occurring proteins. |