Accelerating protein engineering: an integrated framework combines protein language models and epistatic landscape modelingDownload PDF Download PDF Research HighlightOpen accessPublished: 21 August 2026Jeongyoung Koo1,2,Young-Ho Park ORCID: orcid.org/0000-0001-5405-12311,3 &Sun-Uk Kim ORCID: orcid.org/0000-0002-5168-69761,2 Signal Transduction and Targeted Therapy volume 11, Article number: 340 (2026) Cite this articleSave articleView saved researchSubjectsMolecular engineeringIn a recent study published in Science, Tran et al. introduced MULTI-evolve (model guided, universal, targeted installation of multimutants), a machine learning-guided directed evolution (MLDE) framework that rapidly designs hyperactive multimutant proteins.1 By integrating protein language models (PLMs), epistasis-aware modeling, and a high-efficiency multisite mutagenesis platform, the authors address the major combinatorial challenge that has long bottlenecked protein engineering and therapeutic development.Protein function is encoded by amino acid sequence, yet discovering productive mutational combinations remains difficult because protein fitness landscapes are high-dimensional and strongly shaped by epistatic interactions. Directed evolution has long provided a powerful route for protein optimization through repeated mutagenesis and screening.2 More recently, MLDE has expanded accessible sequence space through in silico screening, but its performance still depends on accurately capturing epistatic interactions, the context-dependent effects between mutations that shape protein fitness.3 In particular, many existing approaches require large training datasets, multiple experimental rounds, or labor-intensive synthesis, and remain limited in extrapolating to high-order multimutants because of incomplete epistatic modeling. Tran et al. address these limitations with an end-to-end framework for direct multimutant exploration.MULTI-evolve is built on three conceptual components. First, the framework deploys a PLM zero-shot ensemble approach to nominate function-enhancing single mutations. By combining structure-informed and sequence-based models with z-score normalization, the authors improved the identification of productive mutations compared with individual PLMs alone. Second, they trained fully connected neural networks (FCNNs) on a compact dataset of experimentally characterized single and double mutants to learn epistatic interactions and predict higher-order variants. This data-efficient strategy supported extrapolation from single- and double-mutant data to higher-order combinations, including variants with up to 12 mutations in benchmarking analyses. Third, the authors developed MULTI-assembly, a nicking-based multisite mutagenesis method that supports efficient construction of predicted multimutants across multikilobase coding sequences, thereby reducing dependence on commercial gene synthesis (Fig. 1).Fig. 1Full size imageMULTI-evolve couples PLM priors, epistasis learning, and multisite mutagenesis for single-round multimutant engineering. Starting from a protein of interest, function-enhancing single mutations are identified using PLM zero-shot prediction or existing functional datasets, after which a compact panel of pairwise combinations is experimentally characterized to capture key epistatic interactions. These data train a neural-network model that extrapolates to higher-order combinations and nominates top multimutants, which are then rapidly constructed by MULTI-assembly across multikilobase coding sequences. Application of this framework to APEX, dCasRx, and an anti-CD122 antibody produced synergistic variants with marked gains in catalytic activity, RNA-guided trans-splicing performance, or combined binding-expression properties, establishing MULTI-evolve as a generalizable platform for accelerated protein engineering. Figure created with BioRender.com (https://BioRender.com/6ym60ij)The generalizability of MULTI-evolve was demonstrated across three distinct engineered proteins. For engineered ascorbate peroxidase (APEX), PLM‑guided scanning recapitulated the canonical A134P (APEX2) mutation and identified additional substitutions that yield up to 256‑fold higher activity than wild‑type APEX and up to 4.8‑fold improvement over APEX2 in biochemical and cell‑based assays.4 Notably, even in the absence of A134P, the framework generated variants with up to 13.5-fold higher activity than wild-type APEX, whereas EVOLVEpro achieved only up to 1.9-fold enhancement after five rounds.5 Biochemical and cell-based analyses further showed that the activity gains were not attributable to increased cellular expression.In parallel, MULTI-evolve was applied to engineer the catalytically inactive RNA-targeting CRISPR effector dCasRx (dead Cas13d) for programmable RNA trans-splicing, using deep mutational scanning (DMS)-derived functional data as initialization rather than PLM predictions. This yielded dCasRx multimutants with up to 9.8-fold higher splicing activity and up to 4.5-fold greater trans-splicing efficiency in three endogenous human genes, ITGB1, TFRC, and SMARCA4. The activating mutations E282 and S354P mapped near the crRNA/target RNA duplex, indicating a possible role in improving RNA duplex engagement as the mechanistic basis. These activity gains were also retained, although to a lesser extent, in the context of RNA-guided trans-splicing with Cas editor (RESPLICE), where dCasRx is coupled to a Cas7-11 cis-interfering module, indicating that the engineered variants improve function without additional effector fusions.A clinically relevant application of MULTI-evolve was the multiobjective optimization of HuABC2, a high-affinity anti-CD122 antibody with therapeutic potential for autoimmune diseases such as vitiligo and celiac disease. Antibody development often involves navigating complex Pareto frontiers, where improving one developability parameter can compromise another. In this context, MULTI-evolve identified distinct epistatic architectures underlying expression and binding affinity. By accounting for antagonistic mutational interactions, it achieved a simultaneous 6.5-fold improvement in expression yield and a 2.7-fold enhancement in binding affinity, thereby improving both therapeutic function and manufacturability.Despite these advances, several limitations of MULTI-evolve should be considered. First, its performance depends on the initial identification of a sufficient number of function-enhancing single mutations. The framework may also be difficult to apply when PLM zero-shot predictions do not yield advantageous variants, particularly if evolutionary bias does not correspond to the function of interest and experimental DMS data are lacking. The framework samples sequence space surrounding beneficial mutations, potentially missing evolutionary pathways where optimized variants arise from combinations of mutations that are neutral or deleterious alone. Future iterations could incorporate such mutations to expand the sequence space and capture higher-order positive epistatic effects. Second, its reliance on one-hot sequence encodings in the neural network limits the model’s representational capacity. It does not incorporate broader biophysical features, such as residue-residue contacts or protein structural information, which PLM embeddings can capture and which may enable better structural extrapolation. Moreover, the accuracy of such extrapolation may decline as mutational distance from the training data increases, especially in highly rugged fitness landscapes. Given recent advances in AI-driven protein structure prediction and representation learning, integrating structural embeddings from protein foundation models could improve epistasis prediction and generalization beyond simple sequence encodings.The broader implications of MULTI-evolve lie in its potential to streamline iterative protein engineering workflows. By combining computational prioritization with a compact experimental cycle, MULTI-evolve may reduce the time and experimental burden associated with multimutant optimization. The framework’s capacity for data-efficient, multiobjective optimization is particularly relevant to biopharmaceutical development, where efficacy, manufacturability, and safety often need to be balanced in parallel. As targeted therapeutics increasingly rely on engineered modalities, including antibodies, gene-editing systems, and synthetic enzymes, the ability to design and test synergistic multimutants could expand the practical scope of protein engineering. Collectively, MULTI-evolve provides a data-efficient framework for protein engineering by integrating PLM-guided mutation discovery, epistasis-aware modeling, and combinatorial multisite variant design. This framework may facilitate the translation of computationally designed proteins toward biomedical applications.ReferencesTran, V. Q. et al. Rapid directed evolution guided by protein language models and epistatic interactions. Science 392, eaea1820 (2026).Article CAS PubMed PubMed Central Google Scholar Romero, P. A. & Arnold, F. H. Exploring protein fitness landscapes by directed evolution. Nat. Rev. Mol. Cell Biol. 10, 866–876 (2009).Article CAS PubMed PubMed Central Google Scholar Yang, K. K., Wu, Z. & Arnold, F. H. Machine-learning-guided directed evolution for protein engineering. Nat. Methods 16, 687–694 (2019).Article CAS PubMed Google Scholar Lam, S. S. et al. Directed evolution of APEX2 for electron microscopy and proximity labeling. Nat. Methods 12, 51–54 (2015).Article CAS PubMed Google Scholar Jiang, K. et al. Rapid in silico directed evolution by a protein language model with EVOLVEpro. Science 387, eadr6006 (2025).Article CAS PubMed Google Scholar Download referencesAcknowledgementsThis research was supported by the KRIBB Research Initiative Program (KQM0042611), National Research Foundation (NRF) funded by the Korean government (MSIT) (RS-2021NR057659, RS-2025-00518480, RS-2026-25476968), and the National Research Council of Science & Technology (NST) grant by the Korea government (MSIT) (No. GTL24023-000) and the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (RS-2026-25502637).Author informationAuthors and AffiliationsFuturistic Animal Resource and Research Center, Korea Research Institute of Bioscience and Biotechnology (KRIBB), Cheongju, Republic of KoreaJeongyoung Koo, Young-Ho Park & Sun-Uk KimFunctional genomics Department, KRIBB School, Korea National University of Science and Technology (UST), Daejeon, Republic of KoreaJeongyoung Koo & Sun-Uk KimAdvanced Bioconvergence Department, KRIBB School, Korea National University of Science and Technology (UST), Daejeon, Republic of KoreaYoung-Ho ParkAuthorsJeongyoung KooView author publicationsSearch author on:PubMed Google ScholarYoung-Ho ParkView author publicationsSearch author on:PubMed Google ScholarSun-Uk KimView author publicationsSearch author on:PubMed Google ScholarContributionsJ.K., Y.H.P. and S.U.K. designed, researched, and wrote the manuscript. All authors have read and approved the article.Corresponding authorsCorrespondence to Young-Ho Park or Sun-Uk Kim.Ethics declarationsCompeting interestsThe authors declare no competing interests.Additional informationPublisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.Rights and permissionsOpen Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.Reprints and permissionsAbout this articleDownload PDF