Sequence Characteristics and Codon Bias Analysis of Chloroplast Genomes in Local Population of Suichuan Gougunao Tea Tree

YINMinghua, DONGJiaying, CHENLiyan, LIJunyan, DENGYun, HUANGJia

Chin Agric Sci Bull ›› 2026, Vol. 42 ›› Issue (18) : 150-158.

PDF(2272 KB)
Home Journals Chinese Agricultural Science Bulletin
Chinese Agricultural Science Bulletin

Abbreviation (ISO4): Chin Agric Sci Bull      Editor in chief: Yulong YIN

About  /  Aim & scope  /  Editorial board  /  Indexed  /  Contact  / 
PDF(2272 KB)
Chin Agric Sci Bull ›› 2026, Vol. 42 ›› Issue (18) : 150-158. DOI: 10.11924/j.issn.1000-6850.casb2025-0842

Sequence Characteristics and Codon Bias Analysis of Chloroplast Genomes in Local Population of Suichuan Gougunao Tea Tree

Author information +
History +

Abstract

To analyze the complete chloroplast genome characteristics of the local population of Suichuan Gougunao tea tree, the DNBSEQ-T7 BGI sequencing platform and bioinformatics analysis software were used in this study to analyze the chloroplast genome of the local population of Suichuan Gougunao tea tree. The chloroplast genome of the local population of Suichuan Gougunao tea tree was a typical quadripartite circular structure with a total length of 157099 bp. A total of 135 functional genes were annotated, and 53 simple sequence repeat (SSR) loci and 48 long repeats were detected. The variation range of nucleotide polymorphism (Pi) was 0-0.01265, with weak codon bias. There were a total of 14 optimal codons, with the ndhB gene and rps8 gene undergoing positive selection, while the remaining genes were subject to purifying selection. The local population of Suichuan Gougunao tea tree was closely related to Dehong tea (KJ806279). The divergence time between the local population of Suichuan Gougunao tea tree and Tingpingdong (PP231879) and Liubao tea (OQ281601) was approximately 5.55 million years ago.

Key words

Suichuan Gougunao tea tree / the local population / chloroplast genome / sequence features / codon bias / phylogeny

Cite this article

Download Citations
YIN Minghua , DONG Jiaying , CHEN Liyan , et al . Sequence Characteristics and Codon Bias Analysis of Chloroplast Genomes in Local Population of Suichuan Gougunao Tea Tree[J]. Chinese Agricultural Science Bulletin. 2026, 42(18): 150-158 https://doi.org/10.11924/j.issn.1000-6850.casb2025-0842

References

[1]
肖新华. 江西遂川“狗牯脑”茶叶产业发展的现状及对策[J]. 现代园艺, 2017(3):50-51.
[2]
张贱根, 江新凤, 曾永强, 等. 狗牯脑2号茶树红茶适制性研究[J]. 贵州农业科学, 2024, 52(11):116-123.
[3]
马雪茗, 李海燕, 曾斌, 等. “狗牯脑2号”所制红茶“蜜兰香”关键香气成分分析[J]. 现代食品科技, 2025, 41(3):340-347,349-350.
[4]
吴福茹. 基于代谢组学研究不同产地狗牯脑茶化学成分差异[D]. 南昌: 南昌大学,2024:48-62.
[5]
WANG S, GENG S G, WANG X S, et al. Comparative and phylogenetic analysis of Platycarya longipes and related species based on the complete chloroplast genomes[J]. Genome, 2025, 68:1-12.
[6]
CAI Y F, TIAN M, YANG Y J, et al. Nine complete chloroplast genomes of the Camellia genus provide insights into evolutionary relationships and species differentiation[J]. Scientific reports, 2025, 15(1):8783.
[7]
WANG S, WANG Y H, SUN J H, et al. Comparative chloroplast genome analyses provide new insights into molecular markers for distinguishing Arnebiae Radix and its substitutes (tribe Lithospermeae, Boraginaceae)[J]. Phytomedicine, 2025, 136:156338.
[8]
杨雨青, 谭娟, 汪芳, 等. 茶树叶绿体基因组的研究与应用进展[J]. 生物技术通报, 2024, 40(2):20-30.
多数植物叶绿体基因组为双链环状DNA,具有保守的四分体结构、非重组、单倍体、单亲遗传、进化速率适中、序列和结构高度保守等特征,可为植物进化等提供有用信息。茶树种质资源丰富,其复杂的起源、进化、分类等方面缺乏有效鉴评,而有碍于对其高效保护与创新利用。目前,叶绿体基因组测序技术快速发展并向三代过渡,33份茶树资源的叶绿体基因组已在NCBI公布,茶树叶绿体基因组基因类型、基因序列特征已被部分揭示,且已应用于茶树资源的分类鉴定、起源进化、白化机制等方面研究,但仍有不少茶树资源的叶绿体基因组未被破译。本文简述了植物叶绿体基因组的起源、遗传方式、基本特征;重点述评了茶树叶绿体基因组测序技术、基因组特征、基因类型、基因序列等最新研究成果;概述了茶树叶绿体基因组应用现状;最后探讨了茶树叶绿体基因组未来的应用与发展方向。
[9]
罗祥宗, 胡云飞, 吴淋慧, 等. 茶树叶绿体基因组SNP分子标记的初步研究[J]. 茶叶科学, 2022, 42(6):768-778.
[10]
曾文娟, 刘珊, 文聪, 等. ‘白毫早’叶绿体与线粒体基因组密码子偏好性分析[J]. 茶叶科学, 2025, 45(4):587-603.
[11]
曾文娟, 周伟, 何诗恬, 等. 茶树品种白毫早叶绿体基因组结构特征及其密码子偏好性分析[J]. 江苏农业学报, 2025, 41(7):1398-1411.
[12]
尹明华, 张嘉欣, 乐芸, 等. 茶树大面白叶绿体基因组特征、密码子偏好性及其系统发育分析[J]. 茶叶科学, 2024, 44(3):411-430.
[13]
赵许朋, 崔奎, 耿苗苗, 等. 贵定鸟王茶的叶绿体基因组特征[J]. 西南农业学报, 2023, 36(11):2348-2357.
[14]
佟岩, 黄荟, 王雨华. 森林茶园古茶树大理茶叶绿体基因组密码子偏好性及系统发育研究[J]. 茶叶科学, 2023, 43(3):297-309.
[15]
黎巷汝, 赵雅琦, 张艳, 等. 武夷名丛叶绿体基因组序列特征及系统发育分析[J]. 南方农业学报, 2023, 54(5):1352-1362.
[16]
刘振, 赵洋, 杨培迪, 等. 三倍体茶树‘西莲1号’叶绿体基因组特征及系统发育分析[J]. 茶叶通讯, 2023, 50(2):166-175.
[17]
闫明慧, 刘柯, 王满, 等. 信阳10号叶绿体基因组及其系统进化[J]. 茶叶科学, 2021, 41(6):777-788.
[18]
洪涛. 狗牯脑绿茶多糖的结构特征及体外消化酵解特性[D]. 南昌: 南昌大学,2022:4-5.
[19]
聂永昭, 郭文琼, 梁小金. 狗牯脑夏秋茶速溶茶粉生产工艺探究[J]. 广东蚕业, 2020, 54(12):83-84.
[20]
张贱根, 李琛, 江新凤, 等. 狗牯脑红茶标准化关键加工技术研究[J]. 蚕桑茶叶通讯, 2025(2):15-22.
[21]
高飞, 封新林. 遂川“狗牯脑茶”区域公用品牌建设研究[J]. 农村经济与科技, 2022, 33(22):103-105.
[22]
聂永昭, 郭文琼, 梁小金. 遂川县“狗牯脑”茶产业发展研究[J]. 南方农机, 2021, 52(3):47-48,51.
[23]
黎丽. 遂川县狗牯脑茶生长气候条件分析[J]. 现代农业科技, 2016(11):277.
[24]
袁兴华, 郭小龙, 李源华, 等. 遂川县狗牯脑茶叶生产中主要病虫害发生规律及绿色防控技术集成与推广应用[J]. 福建农业, 2014(8): 161-162.
[25]
RIYAD H, MYLES C, ALASDAIR S, et al. An optimized CTAB method for genomic DNA extraction from green seaweeds (Ulvophyceae)[J]. Applications in plant sciences, 2025, 13(1):1-8.
[26]
GAN Y Y, PING J Y, LIU X J, et al. Repetitive sequences, codon usage bias and phylogenetic analysis of the plastome of Miliusa glochidioides[J]. Biochemical genetics, 2025, 63(4):3329-3346.
[27]
CAMIOLO S, MELITO S, PORCEDDU A. New insights into the interplay between codon bias determinants in plants[J]. DNA research, 2015, 22(6):461-470.
Codon bias is the non-random use of synonymous codons, a phenomenon that has been observed in species as diverse as bacteria, plants and mammals. The preferential use of particular synonymous codons may reflect neutral mechanisms (e.g. mutational bias, G|C-biased gene conversion, genetic drift) and/or selection for mRNA stability, translational efficiency and accuracy. The extent to which these different factors influence codon usage is unknown, so we dissected the contribution of mutational bias and selection towards codon bias in genes from 15 eudicots, 4 monocots and 2 mosses. We analysed the frequency of mononucleotides, dinucleotides and trinucleotides and investigated whether the compositional genomic background could account for the observed codon usage profiles. Neutral forces such as mutational pressure and G|C-biased gene conversion appeared to underlie most of the observed codon bias, although there was also evidence for the selection of optimal translational efficiency and mRNA folding. Our data confirmed the compositional differences between monocots and dicots, with the former featuring in general a lower background compositional bias but a higher overall codon bias. © The Author 2015. Published by Oxford University Press on behalf of Kazusa DNA Research Institute.
[28]
QIAO Z S, LI J Q, ZHANG X L, et al. Genome-wide identification, expression analysis, and subcellular localization of DET2 gene family in Populus yunnanensis[J]. Genes, 2024, 15(2):148.
(1) Background: Brassinosteroids (BRs) are important hormones involved in almost all stages of plant growth and development, and sterol dehydrogenase is a key enzyme involved in BRs biosynthesis. However, the sterol dehydrogenase gene family of Populus yunnanensis Dode (P. yunnanensis) has not been studied. (2) Methods: The PyDET2 (DEETIOLATED2) gene family was identified and analyzed. Three genes were screened based on RNA-seq of the stem tips, and the PyDET2e was further investigated via qRT-PCR (quantitative real-time polymerase chain reaction) and subcellular localization. (3) Results: The 14 DET2 family genes in P. yunnanensis were categorized into four groups, and 10 conserved protein motifs were identified. The gene structure, chromosome distribution, collinearity, and codon preference of all PyDET2 genes in the P. yunnanensis genome were analyzed. The codon preference of this family is towards the A/U ending, which is strongly influenced by natural selection. The PyDET2e gene was expressed at a higher level in September than in July, and it was significantly expressed in stems, stem tips, and leaves. The PyDET2e protein was localized in chloroplasts. (4) Conclusions: The PyDET2e plays an important role in the rapid growth period of P. yunnanensis. This systematic analysis provides a basis for the genome-wide identification of genes related to the brassinolide biosynthesis process in P. yunnanensis, and lays a foundation for the study of the rapid growth mechanism of P. yunnanensis.
[29]
HERSHBERG R, PETROV D A. General rules for optimal codon choice[J]. Plos genetics, 2009, 5(7):e1000556.
[30]
HUH M K. The rates of synonymous and nonsynonymous substitutions in Sorbus aucuparia using nuclear and chloroplast genes[J]. Journal of life science, 2010, 20(4):481-486.
[31]
PENG J, ZHAO Y L, DONG M, et al. Exploring the evolutionary characteristics between cultivated tea and its wild relatives using complete chloroplast genomes[J]. BMC ecology and evolution, 2021, 21:71.
[32]
陈春梅, 马春雷, 马建强, 等. 茶树cpDNA测序及基于cpDNA序列的山茶属植物亲缘关系研究[J]. 茶叶科学, 2014, 34(4):371-380.
[33]
FAN L, LI L, HU Y F, et al. Complete chloroplast genomes of five classical Wuyi tea varieties (Camellia sinensis, Synonym: Thea bohea L.), the most famous Oolong tea in China[J]. Mitochondrial dna part b, 2022, 7(4):655-657.
[34]
钟志敏, 赖小平, 黄松, 等. 基于ycf1的姜科植物条形码鉴定及聚类分析[J]. 中华中医药杂志, 2018, 33(9):4089-4092.
[35]
高娜娜, 赵志礼, 倪梁红. 植物叶绿体ycf15基因应用于药用植物鉴定的前景展望[J]. 中草药, 2017, 48(15):3210-3217.
[36]
KIKUCHI S, ASAKURA Y, IMAI M, et al. A Ycf2-FtsHi heteromeric AAA-ATPase complex is required for chloroplast protein import[J]. Plant cell, 2018, 30(11):2677-2703.
[37]
NELLAEPALLI S, OZAWA S I, KURODA H, et al. The photosystem I assembly apparatus consisting of Ycf3-Y3IP1 and Ycf4 modules[J]. Nature communications, 2018, 9(1):2439.
[38]
YAN C, DU J C, GAO L, et al. The complete chloroplast genome sequence of watercress (Nasturtium officinale R. Br.): genome organization, adaptive evolution and phylogenetic relationships in Cardamineae[J]. Gene, 2019, 699:24-36.
[39]
PARK I, YANG S, CHOI G, et al. The complete chloroplast genome sequences of Aconitum pseudolaeve and Aconitum longecassidatum, and development of molecular markers for distinguishing species in the Aconitum subgenus Lycoctonum[J]. Molecules, 2017, 22(11):2012.
Aconitum pseudolaeve Nakai and Aconitum longecassidatum Nakai, which belong to the Aconitum subgenus Lycoctonum, are distributed in East Asia and Korea. Aconitum species are used in herbal medicine and contain highly toxic components, including aconitine. A. pseudolaeve, an endemic species of Korea, is a commercially valuable material that has been used in the manufacture of cosmetics and perfumes. Although Aconitum species are important plant resources, they have not been extensively studied, and genomic information is limited. Within the subgenus Lycoctonum, which includes A. pseudolaeve and A. longecassidatum, a complete chloroplast (CP) genome is available for only one species, Aconitum barbatum Patrin ex Pers. Therefore, we sequenced the complete CP genomes of two Aconitum species, A. pseudolaeve and A. longecassidatum, which are 155,628 and 155,524 bp in length, respectively. Both genomes have a quadripartite structure consisting of a pair of inverted repeated regions (51,854 and 52,108 bp, respectively) separated by large single-copy (86,683 and 86,466 bp) and small single-copy (17,091 and 16,950 bp) regions similar to those in other Aconitum CP genomes. Both CP genomes consist of 112 unique genes, 78 protein-coding genes, 4 ribosomal RNA (rRNA) genes, and 30 transfer RNA (tRNA) genes. We identified 268 and 277 simple sequence repeats (SSRs) in A. pseudolaeve and A. longecassidatum, respectively. We also identified potential 36 species-specific SSRs, 53 indels, and 62 single-nucleotide polymorphisms (SNPs) between the two CP genomes. Furthermore, a comparison of the three Aconitum CP genomes from the subgenus Lycoctonum revealed highly divergent regions, including trnK-trnQ, ycf1-ndhF, and ycf4-cemA. Based on this finding, we developed indel markers using indel sequences in trnK-trnQ and ycf1-ndhF. A. pseudolaeve, A. longecassidatum, and A. barbatum could be clearly distinguished using the novel indel markers AcoTT (Aconitum trnK-trnQ) and AcoYN (Aconitum ycf1-ndhF). These two new complete CP genomes provide useful genomic information for species identification and evolutionary studies of the Aconitum subgenus Lycoctonum.
[40]
LI W, ZHANG C P, GUO X, et al. Complete chloroplast genome of Camellia japonica genome structures, comparative and phylogenetic analysis[J]. Plos one, 2019, 14(5):e0216645.
[41]
DROUIN G, DAOUD H, XIA J N. Relative rates of synonymous substitutions in the mitochondrial, chloroplast and nuclear genomes of seed plants[J]. Molecular phylogenetics and evolution, 2008, 49(3):827-831.
Previous studies have estimated that, in angiosperms, the synonymous substitution rate of chloroplast genes is three times higher than that of mitochondrial genes and that of nuclear genes is twelve times higher than that of mitochondrial genes. Here we used 12 genes in 27 seed plant species to investigate whether these relative rates of substitutions are common to diverse seed plant groups. We find that the overall relative rate of synonymous substitutions of mitochondrial, chloroplast and nuclear genes of all seed plants is 1:3:10, that these ratios are 1:2:4 in gymnosperms but 1:3:16 in angiosperms and that they go up to 1:3:20 in basal angiosperms. Our results show that the mitochondrial, chloroplast and nuclear genomes of seed plant groups have different synonymous substitutions rates, that these rates are different in different seed plant groups and that gymnosperms have smaller ratios than angiosperms.
[42]
WANG W B, YU H, WANG J H, et al. The complete chloroplast genome sequences of the medicinal plant Forsythia suspensa (Oleaceae)[J]. International journal of molecular sciences, 2017, 18(11):2288.
Forsythia suspensa is an important medicinal plant and traditionally applied for the treatment of inflammation, pyrexia, gonorrhea, diabetes, and so on. However, there is limited sequence and genomic information available for F. suspensa. Here, we produced the complete chloroplast genomes of F. suspensa using Illumina sequencing technology. F. suspensa is the first sequenced member within the genus Forsythia (Oleaceae). The gene order and organization of the chloroplast genome of F. suspensa are similar to other Oleaceae chloroplast genomes. The F. suspensa chloroplast genome is 156,404 bp in length, exhibits a conserved quadripartite structure with a large single-copy (LSC; 87,159 bp) region, and a small single-copy (SSC; 17,811 bp) region interspersed between inverted repeat (IRa/b; 25,717 bp) regions. A total of 114 unique genes were annotated, including 80 protein-coding genes, 30 tRNA, and four rRNA. The low GC content (37.8%) and codon usage bias for A- or T-ending codons may largely affect gene codon usage. Sequence analysis identified a total of 26 forward repeats, 23 palindrome repeats with lengths >30 bp (identity > 90%), and 54 simple sequence repeats (SSRs) with an average rate of 0.35 SSRs/kb. We predicted 52 RNA editing sites in the chloroplast of F. suspensa, all for C-to-U transitions. IR expansion or contraction and the divergent regions were analyzed among several species including the reported F. suspensa in this study. Phylogenetic analysis based on whole-plastome revealed that F. suspensa, as a member of the Oleaceae family, diverged relatively early from Lamiales. This study will contribute to strengthening medicinal resource conservation, molecular phylogenetic, and genetic engineering research investigations of this species.
[43]
HUANG H, SHI C, LIU Y, et al. Thirteen Camellia chloroplast genome sequences determined by high-throughput sequencing: genome structure and phylogenetic relationships[J]. BMC evolutionary biology, 2014, 14:151.
[44]
LI L, HU Y F, HE M, et al. Comparative chloroplast genomes: insights into the evolution of the chloroplast genome of Camellia sinensis and the phylogeny of Camellia[J]. BMC genomics, 2021, 22(1):138.
\n Chloroplast genome resources can provide useful information for the evolution of plant species. Tea plant (\n Camellia sinensis\n ) is among the most economically valuable member of\n Camellia\n. Here, we determined the chloroplast genome of the first natural triploid Chinary type tea (‘Wuyi narcissus’ cultivar of\n Camellia sinensis var. sinensis\n,\n CWN\n ) and conducted the genome comparison with the diploid Chinary type tea (\n Camellia sinensis var. sinensis\n,\n CSS\n ) and two types of diploid Assamica type teas (\n Camellia sinensis var. assamica\n : Chinese Assamica type tea,\n CSA\n and Indian Assamica type tea,\n CIA\n ). Further, the evolutionary mechanism of the chloroplast genome of\n Camellia sinensis\n and the relationships of\n Camellia\n species based on chloroplast genome were discussed.\n
[45]
CAMPBELL W H, GOWRI G. Codon usage in higher plants, green algae, and cyanobacteria[J]. Plant physiology, 1990, 92(1):1-11.
Codon usage is the selective and nonrandom use of synonymous codons by an organism to encode the amino acids in the genes for its proteins. During the last few years, a large number of plant genes have been cloned and sequenced, which now permits a meaningful comparison of codon usage in higher plants, algae, and cyanobacteria. For the nuclear and organellar genes of these organisms, a small set of preferred codons are used for encoding proteins. Codon usage is different for each genome type with the variation mainly occurring in choices between codons ending in cytidine (C) or guanosine (G) versus those ending in adenosine (A) or uridine (U). For organellar genomes, chloroplastic and mitochrondrial proteins are encoded mainly with codons ending in A or U. In most cyanobacteria and the nuclei of green algae, proteins are encoded preferentially with codons ending in C or G. Although only a few nuclear genes of higher plants have been sequenced, a clear distinction between Magnoliopsida (dicot) and Liliopsida (monocot) codon usage is evident. Dicot genes use a set of 44 preferred codons with a slight preference for codons ending in A or U. Monocot codon usage is more restricted with an average of 38 codons preferred, which are predominantly those ending in C or G. But two classes of genes can be recognized in monocots. One set of monocot genes uses codons similar to those in dicots, while the other genes are highly biased toward codons ending in C or G with a pattern similar to nuclear genes of green algae. Codon usage is discussed in relation to evolution of plants and prospects for intergenic transfer of particular genes.
[46]
WANG L J, ROOSSINCK M J. Comparative analysis of expressed sequences reveals a conserved pattern of optimal codon usage in plants[J]. Plant molecular biology, 2006, 61(4/5):699-710.
[47]
SALIM H M W, CAVALCANTI A R O. Factors influencing codon usage bias in genomes[J]. Journal of the brazilian chemical society, 2008, 19(2):257-262.
[48]
HERSHBERG R, PETROV D A. Selection on codon bias. TL-42[J]. Annual review of genetics, 2008, 42(1):287-299.
[49]
NOCK C J, WATERS D L E, EDWARDS M A, et al. Chloroplast genome sequences from total DNA for plant identification[J]. Plant biotechnology journal, 2015, 9(3):328-333.
PDF(2272 KB)

Accesses

Citation

Detail

Sections
Recommended

/

〈 〉