research

Our overarching vision is to create an autonomous, AI-driven ecosystem for biomolecular engineering, to address key challenges such as sustainable synthesis, therapeutics, environmental remediation, and more:

Research overview

To this end, we develop data-driven AI methods to engineer enzymes and biomolecules with novel and useful functions, and we integrate these with real-world experimental workflows. Our research sits at the intersection of machine learning, protein engineering, computational biology, and computational chemistry. We focus on three central themes:

  • data foundations: how to collect and utilize relevant empirical data and physics-based models?
  • modeling: how to develop efficient AI methods for multi-modal tasks?
  • application: how to bridge computational methods with real-world experimental constraints?

Research Thrusts

Creating a computational platform for functional biomolecule discovery

Research image 1
Less than 1% of known protein sequences are annotated with functions, so an improved method to discover useful enzymes with novel functions is needed. Existing bioinformatics (e.g., BLAST) and ML methods can automatically retrieve unannotated sequences from databases (given a desired function) but struggle for sequences with remote homology to known proteins and previously uncharacterized functions. We are building new enzyme sequence-function datasets as the basis for developing a multi-modal foundation model that captures a mapping from enzymes to their functions. Our goal is to capture the universe of enzymatic catalysis, to enable enzyme annotation and retrieval for applications such as metabolic engineering for natural product synthesis, enzymes for biomass upcycling, and novel gene editors for therapeutics.

Related work:

Reimagining deep learning methods for structure-based biomolecular design

Research image 2
We aim to venture beyond the existing paradigm of de novo protein design, which largely simplifies proteins as static, structured entities and lacks physical consideration of dynamics and electronics, including when modeling protein interactions with other biomolecules. Many functions cannot be accessed with such existing approaches, so we are developing a generative model for protein structure where design is conditioned on desired physical interactions with substrate(s) to enable specific allosteric modulation or enzymatic chemistry. To facilitate to this effort, we improve molecular models of biocatalysis by integrating AI with approaches from electronic structure theory and molecular dynamics.

Related work:

Accelerating biomolecular optimization with generative modeling

Research image 3
After discovering or designing a protein with some level of desired function, proteins often need to be optimized (modified) to maximize fitness for specific objectives such as stability, activity, affinity, etc. We develop methods that are more efficient than directed evolution, through ML-assisted optimization, which involves an iterative cycle of collecting labeled data through expensive wet-lab measurements and using these data to update an ML model and suggest new sequences to further explore, known as active learning. To this end, we develop state-of-the-art ML methods that use labeled data to shift the distribution of generative models to conditionally sample sequences with higher fitness. Ulimately, we aim to integrate these strategies into real-world active learning workflows that are increasingly powered by agentic AI and autonomous discovery.

Related work:

If you are interested in joining the group, please get in touch!