AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (5.9 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Leveraging pre-trained AI models for robust promoter sequence design in synthetic biology

Gui Yang1,#Yijie Chen1,#Qinghua Guo1,#Xinyang Li1Zhen Zhou1,2( )
Nanjing Drum Tower Hospital Center of Molecular Diagnostic and Therapy, State Key Laboratory of Pharmaceutical Biotechnology, Jiangsu Engineering Research Center for MicroRNA Biology and Biotechnology, NJU Advanced Institute of Life Sciences (NAILS), School of Life Sciences, Nanjing University, Nanjing 210023, China
Institute of Artificial Intelligence Biomedicine, Nanjing University, Nanjing 210023, China

#Gui Yang, Yijie Chen and Qinghua Guo contributed equally to this work.

Show Author Information

Abstract

Although artificial intelligence (AI) has begun to be applied in synthetic biology, it is limited by its reliance on large amounts of high-quality data, which presents a significant challenge in synthetic biology. Pre-trained models have profoundly influenced natural language processing by enabling systems to understand and generate human language with remarkable accuracy and efficiency by capturing complex linguistic patterns and contextual nuances. This study applies the concept of pre-trained models to promoter sequence analysis through an innovative pre-training and fine-tuning paradigm. Our analysis reveals that pre-trained DNA models, particularly DNABERT, consistently outperform non-pre-trained models in predicting promoter expression levels across various dataset sizes. Building on DNABERT's strengths, we developed the AI model Pymaker, which specializes in predicting yeast promoter expression levels. Additionally, we introduced a novel base mutation model to simulate promoter mutations, enabling the generation of new promoter sequences. By integrating Pymaker with this mutation model, we effectively screened for high-expression, mutation-resistant promoters. Experimental validation in Saccharomyces cerevisiae showed that these selected promoters significantly enhanced LTB protein expression. Notably, Pymaker’s predictions demonstrated superior accuracy, achieving a three-fold increase in protein expression compared to traditional promoters. Our findings highlight the potential of Pymaker not only to identify robust promoters but also to significantly reduce reliance on conventional, labor-intensive experimental methods, heralding a new era in synthetic biology and genetic engineering with practical applications in biopharmaceuticals.

Graphical Abstract

Electronic Supplementary Material

Download File(s)
Supplementary Materials.pdf (179.7 KB)

References

【1】
【1】
 
 
Biophysics Reports
Pages 126-135

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Yang G, Chen Y, Guo Q, et al. Leveraging pre-trained AI models for robust promoter sequence design in synthetic biology. Biophysics Reports, 2026, 12(2): 126-135. https://doi.org/10.52601/bpr.2025.240033

165

Views

5

Downloads

2

Crossref

2

Scopus

0

CSCD

Received: 23 July 2024
Accepted: 15 January 2025
Published: 30 April 2026
© The Author(s) 2026

This article is licensed under a Creative Commons Attribution 4.0 International (CC BY 4.0) License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.