Perspective

Machines That Write DNA

By Kyle Harrison

Updated

March 8, 2025

Reading Time

4 min

Last May, we published a deep dive on how AI is breathing new life into biological research. And today, the field of biology is standing on the precipice of that transformation. In February 2025, Arc Institute unveiled Evo 2, an AI model trained on over 100K species' DNA, capable of identifying disease-causing mutations and generating entirely new genomes. This breakthrough represents a fundamental shift from merely deciphering life's code to actively rewriting it from scratch.

Evo 2 was trained on 9.3 trillion nucleotides extracted from more than 128K complete genomes across bacteria, archaea, phages, humans, plants, and various eukaryotes. In tests with the BRCA1 gene, it achieved over 90% accuracy in predicting which variants might cause breast cancer.

"Our development of Evo 1 and Evo 2 represents a key moment in the emerging field of generative biology, as the models have enabled machines to read, write, and think in the language of nucleotides," explains Patrick Hsu, Arc Institute Co-Founder and Core Investigator.

The journey to decode DNA began in 1869 when Friedrich Miescher first isolated what he called "nuclein" from cells. By 1953, Watson and Crick had revealed DNA's double helix structure, and by the 1970s, scientists were sequencing complete genes. These discoveries laid the foundation for genetic engineering and eventually the Human Genome Project, which fully mapped our genetic blueprint by 2003.

DNA analysis and computing has a rich history dating to the 1960s when Margaret Oakley Dayhoff pioneered using computers for biology research. This computational approach accelerated dramatically with next-generation sequencing technologies, allowing scientists to generate massive genomic datasets.

Each breakthrough in understanding genetics sparked new questions about function, regulation, and manipulation. The tension between unraveling nature's complexity and developing tools to modify it has driven progress, influenced by ethical considerations, computing advances, and funding priorities.

Evo 2 represents a quantum leap in technical capability. The model can process genetic sequences of up to 1 million nucleotides at once, enabling it to understand relationships between distant parts of a genome. This eight-fold increase over its predecessor was enabled by an architecture called StripedHyena 2, developed with assistance from OpenAI's Greg Brockman.

Evo2 offers immediate and tangible benefits for researchers and clinicians. The model excels at identifying disease-causing mutations in human genes with over 90% accuracy, potentially eliminating countless hours of expensive lab work while accelerating both diagnosis and drug development.

Researchers can use Evo2 to design gene therapies with cell-specific targeting, reducing side effects. As co-author Hani Goodarzi explains: "If you have a gene therapy that you want to turn on only in neurons to avoid side effects, or only in liver cells, you could design a genetic element that is only accessible in those specific cells."

The Arc Institute has prioritized accessibility by providing three user interfaces: a developer-friendly GitHub repository, integration with NVIDIA's BioNeMo framework to accelerate scientific research, and Evo Designer—a user-friendly interface requiring minimal technical expertise.

"We think of this as enabling an app store for biology," says Patrick Hsu, emphasizing that Evo2 functions as a foundation layer upon which specialized applications can be built. The fully open-source approach allows researchers to fine-tune the model for specific genes or organisms.

The significance extends beyond current capabilities. As Chief Technology Officer Dave Burke notes, researchers will likely discover "beneficial uses for Evo2 we haven't even imagined yet," from predicting how mutations affect protein function to designing genetic elements with novel properties. These innovations could revolutionize everything from disease treatment to sustainable bio-manufacturing.

As Evo models continue to scale, they may soon analyze a person's entire genome to predict disease risks, design targeted therapies, and perhaps craft biological structures never before seen in nature. The path from reading life's code to writing it has only just begun.

Important Disclosures

This material has been distributed solely for informational and educational purposes only and is not a solicitation or an offer to buy any security or to participate in any trading strategy. All material presented is compiled from sources believed to be reliable, but accuracy, adequacy, or completeness cannot be guaranteed, and Contrary LLC (Contrary LLC, together with its affiliates, “Contrary”) makes no representation as to its accuracy, adequacy, or completeness.

The information herein is based on Contrary beliefs, as well as certain assumptions regarding future events based on information available to Contrary on a formal and informal basis as of the date of this publication. The material may include projections or other forward-looking statements regarding future events, targets or expectations. Past performance of a company is no guarantee of future results. There is no guarantee that any opinions, forecasts, projections, risk assumptions, or commentary discussed herein will be realized. Actual experience may not reflect all of these opinions, forecasts, projections, risk assumptions, or commentary.

Contrary shall have no responsibility for: (i) determining that any opinions, forecasts, projections, risk assumptions, or commentary discussed herein is suitable for any particular reader; (ii) monitoring whether any opinions, forecasts, projections, risk assumptions, or commentary discussed herein continues to be suitable for any reader; or (iii) tailoring any opinions, forecasts, projections, risk assumptions, or commentary discussed herein to any particular reader’s objectives, guidelines, or restrictions. Receipt of this material does not, by itself, imply that Contrary has an advisory agreement, oral or otherwise, with any reader.

Contrary is registered with the Securities and Exchange Commission as an investment adviser under the Investment Advisers Act of 1940. The registration of Contrary in no way implies a certain level of skill or expertise or that the SEC has endorsed Contrary. Investment decisions for Contrary clients are made by Contrary. Please note that, although Contrary manages assets on behalf of Contrary clients, Contrary clients may take any position (whether positive or negative) with respect to the company described in this material. The information provided in this material does not represent any investment strategy that Contrary manages on behalf of, or recommends to, its clients.

Different types of investments involve varying degrees of risk, and there can be no assurance that the future performance of any specific investment, investment strategy, company or product made reference to directly or indirectly in this material, will be profitable, equal any corresponding indicated performance level(s), or be suitable for your portfolio. Due to rapidly changing market conditions and the complexity of investment decisions, supplemental information and other sources may be required to make informed investment decisions based on your individual investment objectives and suitability specifications. All expressions of opinions are subject to change without notice. Investors should seek financial advice regarding the appropriateness of investing in any security of the company discussed in this presentation.

Please see www.contrary.com/legal for additional important information.

Authors

Kyle Harrison

General Partner @ Contrary

Kyle leads Contrary’s investing efforts for companies from seed to scale. He’s previously worked at firms like Index and Coatue investing in companies like Databricks, Snowflake, Snyk, Plaid, Toast, and Persona.

See articles

© 2026 Contrary Research · All rights reserved

Privacy Policy

By navigating this website you agree to our privacy policy.