
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

39231996
3830
10.1038/s41597-024-03830-y
Data Descriptor
A near-complete chromosome-level genome assembly of looseleaf lettuce (Lactuca sativa var. crispa)
Zhang Bin 12
Xue Yingfei 3
http://orcid.org/0000-0002-7600-5647
Liu Xue 124
Ding Haifeng 15
Yang Yesheng 5
Wang Chenchen 6
Xu Zhaoyang 6
http://orcid.org/0000-0003-3131-7804
Zhou Jun 6
Sun Cheng 3
Tang Jinfu jftang@icmm.ac.cn

4
Li Dayong lidayong@nercv.org

1278
1 https://ror.org/04trzn023 grid.418260.9 0000 0004 0646 9053 National Engineering Research Center for Vegetables, Beijing Vegetable Research Center, Beijing Academy of Agriculture and Forestry Science, Beijing, 100097 P. R. China
2 https://ror.org/04trzn023 grid.418260.9 0000 0004 0646 9053 State Key Laboratory of Vegetable Biobreeding, Beijing Vegetable Research Center, Beijing Academy of Agriculture and Forestry Science, Beijing, 100097 P. R. China
3 https://ror.org/005edt527 grid.253663.7 0000 0004 0368 505X College of Life Sciences, Capital Normal University, Beijing, 100048 P. R. China
4 https://ror.org/042pgcv68 grid.410318.f 0000 0004 0632 3409 State Key Laboratory for Quality Ensurance and Sustainable Use of Dao-di Herbs, National Resource Center for Chinese Materia Medica, China Academy of Chinese Medical Sciences, Beijing, 100700 P. R. China
5 Jingyan Yinong (Beijing) Seed Sci-Tech Co., Ltd., Beijing, 100097 P. R. China
6 https://ror.org/01wy3h363 grid.410585.d 0000 0001 0495 1805 College of Life Sciences, Shandong Normal University, Jinan, 250014 P. R. China
7 Beijing Key Laboratory of Vegetable Germplasms Improvement, Beijing, 100097 P. R. China
8 Key Laboratory of Biology and Genetics Improvement of Horticultural Crops (North China), Beijing, 100097 P. R. China
4 9 2024
4 9 2024
2024
11 96114 3 2024
27 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Lettuce (Lactuca sativa L., Asteraceae) is one of the most important vegetable crops, known for its various horticultural types and significant morphological variation. The first reference genome of lettuce, a crisphead type (L. sativa var. capitata cv. Salinas), was previously released. Here, we reported a near-complete chromosome-level reference genome for looseleaf lettuce (L. sativa var. crispa). PacBio high-fidelity sequencing, Oxford Nanopore, and Hi-C technologies were employed to produce genome assembly. The final assembly is 2.59 Gb in length with a contig N50 of 205.47 Mb, anchored onto nine chromosomes, containing 14 recognizable telomeres and only 11 gaps. Repetitive sequences account for 77.11% of the genome, and 41,375 protein-coding genes were predicted, with 99.10% of these assigned functional annotations. This chromosome-level genome enriched genomic resources for various horticultural types of lettuce and will facilitate the characterization of morphological variation and genetic improvement in lettuce.

Subject terms

Genome informatics
Plant sciences
Beijing Joint Research Program for Germplasm Innovation and New Variety Breeding (G20220628003-01), Key Project at Central Government Level: the Ability Establishment of Sustainable Use for Valuable Chinese Medicine Resources (2060302), Beijing Innovation Consortium of Agriculture Research System (BAIC01-2023), Innovation and Development Program of Beijing Vegetable Research Center (KYCX202304), and Collaborative Innovation Program of Beijing Vegetable Research Center (XTCX202302).issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

Lactuca sativa L. (Asteraceae), known as lettuce, is considered one of the most important vegetable crops1–5. Originating in the coastal Mediterranean regions, lettuce was featured in Egyptian tombs around 2,500 BC2,3. Today, lettuce is cultivated as diverse horticultural varieties for different purposes, including leafy types (looseleaf, crisphead, romaine, and butterhead) and non-leafy types (stem and oilseed), each with distinct morphological characteristics6,7. Leafy lettuces, particularly looseleaf and crisphead, are consumed globally in salads and hamburgers, and are also popular in hotpot cuisine in China and grilled with red meat in other parts of Asia. Looseleaf lettuce, compared to crisphead, grows faster, can be harvested earlier, and has better tolerates to abiotic stress. Thus, looseleaf lettuce is an important horticultural type for the annual leafy vegetable supply, and genomic research could greatly enhance its economic value.

A high-quality reference genome is crucial for identifying genetic variations, conducting phylogenetic research, and facilitating molecular marker-assisted breeding. As a representative species of the genus Lactuca in the Asteraceae family, the first reference genome for a crisphead lettuce type (L. sativa var. capitata cv. Salinas) was released in 2017, with a genome size of 2.38 Gb and contig N50 of 36 Kb8. With advancements in sequencing technology and broader use of Lactuca species, additional genome assemblies have been published, including those for two wild relatives (L. saligna and L. virosa), and one stem lettuce (L. sativa var. angustana cv. Yanling1)9–11. Although these data are useful for identifying intraspecific variation, only two chromosome-level genome assemblies of cultivated lettuce (the crisphead and stem types) have been generated to date8,11. A single or limited number of reference genomes for an economically important crop is insufficient for exploring genetic diversity, which hinders genomic research and molecular breeding12,13. A high-quality genome assembly for the looseleaf type is crucial for identifying genetic variations, inferring phylogenetic relationships among different horticultural types, and facilitating comparative genomic analysis and genetic improvement in lettuce.

In this study, we generated a chromosome-level and near-complete reference genome assembly for looseleaf lettuce (L. sativa var. crispa cv. Green Elegance) using PacBio high-fidelity reads (~46×), Oxford Nanopore reads (~13×), Illumina short reads (~50.39×), and Hi-C reads (~97×). The assembled genome (Green Elegance) had a total length of 2.59 Gb, with a contig N50 of 205.47 Mb and a BUSCO completeness score of 98.39%. A total of 2,580.61 Mb (99.61%) of the genome sequences were anchored to nine chromosomes, featuring 14 recognizable telomeres and 11 gaps. Genome annotation predicted 41,375 protein-coding genes and 77.11% repetitive sequences. These genomic resources provide a roadmap for further genetic and evolutionary investigation.

Methods

Sample collection, library construction and sequencing

Looseleaf lettuce (Lactuca sativa var. crispa cv. Green Elegance) was provided by the Beijing Vegetable Research Center, Beijing Academy of Agriculture and Forestry Science, Beijing, China (Fig. 1). The seedlings were grown in a growth chamber at the Beijing Vegetable Research Center under a photoperiod of 16-hour light (200 μmol m−2 s−1) and 8-hour dark at 25 °C. Fresh and healthy leaves were collected at the rosette stage and immediately frozen in liquid nitrogen for genome survey and sequencing (Table 1). For transcriptomic sequencing, samples included mature leaves, young seedlings (including roots), and inflorescence (Table 1). Newly developed tender leaves, maintained under moist and low-temperature conditions, were used to construct the Hi-C library (Table 1).Fig. 1 Morphological characteristics of the looseleaf lettuce (Lactuca sativa var. crispa) cultivar Green Elegance sequenced used for genome assembly. Images show the Looseleaf lettuce (cv. Green Elegance) in the field (a) and after harvest (b), including a full-expanded leaf, longitudinal section, the top view and side view of the Green Elegance plant.

Table 1 Summary of the sequencing data generated for the looseleaf lettuce (L. sativa var. crispa cv. Green Elegance) genome assembly.

Library types	Type	Sample	Molecular	Platform	Insert size	Data size (Gb)	Coverage	Application	
Unmap/Short-Read	PE1	Leaf	DNA	Illumina Novaseq 6000	350 bp	124.16 Gb	50.39×	Survey	
SMRTbell	CCS2	Leaf	DNA	PacBio Sequel II	17.26 Kb	109.27 Gb	46×	Assembly	
Hi-C	PE	Leaf	DNA	Illumine Novaseq 6000	300–700 bp	251.37 Gb	97×	Pseudo-chromosome construction	
Direct-cDNA	PE	Leaf	RNA	Illumina Novaseq 6000	350 bp	11.92 Gb	5×	Annotation	
Direct-cDNA	PE	Seedling	
Direct-cDNA	PE	Inflorescence	
ONT	Single	Leaf	DNA	PromethION	100 Kb	32.55 Gb	13×	Gap filling	
1CCS: Circular Consensus Sequencing.

2PE, Paired-end.

High molecular weight genomic DNA was extracted from leaves using a modified CTAB (cetyltrimethylammonium bromide) method14. RNA was removed by adding RNase A. The quality of the DNA was assessed using agarose gel electrophoresis, which confirmed excellent integrity of the DNA molecules.

For Illumina sequencing, a short-read library with an average insert size of 350 bp was constructed and sequenced on an Illumina Novaseq platform (Illumina, CA, USA) using the PE150 program. This yielded 135.8 Gb of raw data. Finally, 124.16 Gb (50.39×) of clean reads were obtained for genome size estimation, sequence correction, and assessment of heterozygosity and repeat content (Table 1 and Fig. 2).Fig. 2 K-mer analysis of L. sativa var. crispa cv. Green Elegance. The plot displays the frequency of 21 K-mers, with the genome size of L. sativa cv. Green Elegance estimated at 2.46 Gb, a heterozygosity rate of 0.211%, and a duplication rate of 1.66%.

For PacBio HiFi sequencing, genomic DNA was fragmented to ~15 Kb to construct a long-read library following the manufacturer’s instructions (Pacific Biosciences, CA, USA). The library was sequenced on a PacBio Sequel II platform using Circular Consensus Sequencing (CCS) mode. The SMRTbell library was constructed using the SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences). Library size and quantity were assessed using the FEMTO Pulse (Agilent Technologies, Wilmington, DE) and the Qubit dsDNA HS Assay Kit (Life Technologies, Carlsbad, CA, USA). The library was loaded at a concentration of 55 pM using diffusion loading. Single-molecule real-time (SMRT) sequencing was conducted on a single 8 M SMRT Cell on the Sequel II System. After filtering out the low-quality reads and sequence adapters, we obtained 109.27 Gb (46×) of clean subreads with a reads-length N50 of 17.81 Kb.

Long-read sequencing using the PromethION platform from Oxford Nanopore Technologies (ONT) was performed to fill assembly gaps. High-quality genomic DNA was fragmented to ~8 Kb using a gTube, and the library was constructed with the Ligation Sequencing Kit 1D (Nanopore, SQK-LSK109). We generated 85.80 Gb of raw data, and after filtering out adapters, low-quality reads, and reads shorter than 2 Kb reads, 32.55 Gb clean reads with a clean N50 length of 100,148 Kb were obtained (Table 1).

Genome size and heterozygosity estimation

Short Illumina reads were quality-filtered using fastp15 (v0.12.4; settings ‘-q 10 -u 50 -y -g -Y 10 -e 20 -l 100 -b 150 -B 150’). The quality-filtered reads were used for genome size estimation. We counted the 21-kmers with Jellyfish16 (v2.1.4; k-mer size 21), and Genomescope17 (v2.0; default settings) were used to estimate a genome size of 2.46 Gb, and a genome-wide heterozygosity rate of 0.21% of sites (Fig. 2).

De novo genome assembly

High-accuracy Circular Consensus Sequencing (CCS) data were used to generate 156 contigs, with the longest contig length of 282.47 Mb and an N50 length of 205.47 Mb using hifiasm (v 0.16) software18 (Table 2). This resulted in a total genome sequence size of 2.59 Gb.Table 2 Comparison of genome assemblies of L. sativa var. crispa, L. sativa var. capitata, L. sativa var. angustana, L. saligna, and L. virosa.

Feature	L. sativa var. crispa (this study)	L. sativa var. capitata8	L. sativa var. angustana11	L. saligna9	L. virosa10	
Number of contigs	156	477	2053	43,198	136,206	
Contig N50	205.47 Mb	12.52 Mb	4.71 Mb	86.84 Kb	135.89 Kb	
Longest contig	282.47 Mb	75.54 Mb	20.98 Mb	794 Kb	1.02 Mb	
Total size of assembled contigs	2.59 Gb	2.59 Gb	2.60 Gb	2.03 Gb	2.51 Gb	
Number of scaffolds	77	93	649	10	5,855	
Scaffold N50	320.76 Mb	325 Mb	332.3 Mb	238.63 Mb	316.91 Mb	
Longest scaffold	407.16 Mb	410.30 Mb	397.86 Mb	420.74 Mb	422.78 Mb	
Total size of assembled scaffold	2.59 Gb	2.59 Gb	2.62 Gb	2.22 Gb	3.45 Gb	
GC content	37.96%	37.92%	37.91%	37.68%	38.00%	
Gap number	11	384	1,404	43,188	130,351	

To anchor contigs, 251.37 Gb of clean reads pairs from the Hi-C library were mapped to the polished Green Elegance genome using BWA (bwa-0.7.17) with the default parameters. Invalid reads, such as self-ligation, non-ligation, PCR amplification, and random breaks, were filtered out. After correction and filtration, we obtained 77 high-accuracy scaffolds with a scaffold N50 length of 320.76 Mb and a total scaffold length of 2,590.68 Mb (Table 2). We successfully anchored 2,590.61 Mb (100%) of the genome into nine groups, which were designated as nine chromosomes of Green Elegance, using the agglomerative hierarchical clustering method in Lachesis19 (Fig. 3). Lachesis was then used to order and orient the clustered contigs. A total of 2,580.61 Mb (99.61%) was successfully ordered and oriented on the nine chromosomes (Table 3). The Hi-C contact heatmap, generated using Hicexplorer v3.720, revealed nine distinct groups based on interaction intensities between bins (a bin size of 800 Kb), indicating high quality of chromosome construction (Fig. 4). The final chromosomal-level assembly had chromosomal lengths ranging from 205,466,188 bp to 407,155,607 bp, encompassing 99.6% of the total sequence (Table 3). After gap filling with ONT sequencing data, 11 gaps remained across eight chromosomes, with one chromosome being complete. Fourteen telomeres, including 11 complete telomeres longer than 1 Kb, were distributed across the nine chromosomes (Fig. 5, Tables 2 and 3). This genome assembly of L. sativa var. crispa cv. Green Elegance represents a significant improvement in genome continuity (contig N50), gap number, and chromosome anchoring compared to the other sequenced Lactuca plants, including L. sativa var. capitata cv. Salinas, L. sativa var. angustana cv. Yanling1, L. saligna, and L. virosa (Table 2).Fig. 3 Circos plot of L. sativa var. crispa cv. Green Elegance genome assembly. Chromosome ideograms (a), transposable element (TE) density (b), simple sequence repeat (SSR) density (c), gene density (d), and GC content (e) were shown from the outermost to the innermost layer.

Table 3 Statistics of the L. sativa var. crispa chromosomes after assembly by Hi-C and gap filling.

Group	Group Length (bp)	Telomere length	Gap number	Order Length (bp)	
Chr 1	255,248,828	20.82 Kb	2	255,189,194	
Chr 2	236,310,553	23.82 Kb/8.42 Kb	2	236,257,685	
Chr 3	320,756,050	553 bp/8.28 Kb	1	320,756,050	
Chr 4	407,188,151	12.57 Kb/10.40 Kb	1	407,155,607	
Chr 5	369,975,299	711 bp/17.97 Kb	1	369,975,299	
Chr 6	205,493,475	—	0	205,466,188	
Chr 7	213,864,026	271 bp/ 1.08 Kb	1	213,762,566	
Chr 8	352,326,448	12.75 Kb	1	342,591,950	
Chr 9	229,451,860	10.29 Kb/ 17.89 Kb	2	229,451,860	
Total (Ratio %)	2,590,614,690		11	2,580,606,399 (99.61%)	

Fig. 4 Hi-C heatmap of L. sativa var. crispa cv. Green Elegance. The intra-chromosome interaction density is represented by varying colors and calculated using a bin size of 800 K.

Fig. 5 The distribution of telomeres and genomic gaps on nine chromosomes of L. sativa var. crispa cv. Green Elegance. After assembly and gap filling, there are 9 centromeres, 14 telomeres and 11 gaps across the nine chromosomes. Gene density (bin size = 1 Mb) is also shown.

Repetitive sequences annotation

Transposon elements (TE) were identified using a combination of homology-based and de novo approaches. A de novo repeat library was first constructed with RepeatModeler (http://www.repeatmasker.org/RepeatModeler/)21. Full-length long terminal repeat retrotransposons (FL-LTR-RTs) were identified using LTRharvest (v1.5.9)22 and LTR_finder (v2.8)23, and a high-quality library was produced with LTR_retriever24. The de novo TE sequences library and known TE sequences from Dfam (v3.5) database were combined to create the final TE sequence set for the Green Elegance genome, which was classified using RepeatMasker (v4.12)25. Tandem repeats were annotated using Tandem Repeats Finder (TRF 409)26 and the MIcroSAtellite identification tool (MISA v2.1)27 with the default parameters (definition: 1–10 2–6 3–5 4–5 5–5 6-5; interruptions: 100). In total, transposon elements and tandem repeats accounted for 77.11% and 4.14% of the Green Elegance genome sequence, respectively, amounting to 2.00 Gb and 107.58 Mb (Table 4).Table 4 Statistics of repetitive element annotation.

Type	Number of elements	Sequence length (bp)	Percent of genome (%)	
Transposable elements	Retroelement	LTR element	Copia	388,930	529,508,339	20.44	
ERV	8,669	1,408,069	0.05	
Gypsy	584,022	733,625,108	28.32	
Ngaro	493	29,296	0	
Pao	921	518,697	0.02	
Non-LTR elements	SINE	4,634	558,319	0.02	
LINE	39,143	14,543,128	0.56	
Retroelement total	1,746,533	1,730,501,429	66.8	
DNA transposon	582,444	267,088,074	10.31	
Total of retroelement and DNA transposon	2,328,977	1,997,589,503	77.11	
Tandem Repeats	Microsatellite (1–9 bp units)	712,396	24,728,216	0.95	
Minisatellite (10–99 bp units)	615,963	61,419,180	2.37	
Satellite (>= 100 bp units)	41,600	21,223,020	0.82	
Total	1,369,959	107,370,416	4.14	

Gene prediction and functional annotation of protein-coding genes

Three approaches—de novo prediction, homology search, and transcript-based assembly—were integrated for annotating protein-coding genes in the genome (Table 5). De novo gene models were predicted using two ab initio gene-prediction software tools, Augustus (v3.1.0)28 and SNAP (Korf, 2004). For homolog-based prediction, GeMoMa (v1.7) was used with reference gene models from the various species, including Arabidopsis thaliana, Oryza sativa, L. sativa var. capitata cv. Salinas, L. sativa var. angustana, L. serriola, L. virosa, Helianthus annus, Taraxacun kok-saghyz, and Artemisia annua. For transcript-based prediction, RNA-sequencing data were mapped to reference genome using Hisat (v2.1.0)29 and assembled with Stringtie (v 2.1.4)17. GeneMarkS-T (v5.1) was used to predict genes based on these assembled transcripts. Additionally, PASA (v2.4.1) was employed to predict genes based on unigenes and full-length transcripts from PacBio/ONT sequencing assembled by Trinity (v2.11)30. Gene models from these approaches were integrated using EVM (v1.1.1) and updated with PASA. In total, 41,375 protein-coding genes with an average length of 3,744 bp were predicted in the Green Elegance genome (Table 6).Table 5 Statistics of gene number by different annotation methods.

Method	Software	Species	Gene number	
Ab initio	Augustus	—	66,681	
	SNAP	—	174,141	
Homology-based	GeMoMa	A. annua	43,480	
		A. thaliana	27,638	
		H. annuus	48,387	
		L. saligna	59,066	
		L. sativa_Salinas_v11	56,505	
		L. sativa_angustana	40,589	
		L. virosa	54,179	
		O. sativa	38,547	
		T. kok-saghyz	50,313	
RNAseq	GeneMarkS-T	—	21,439	
	PASA	—	22,357	

Table 6 Summary of gene annotation.

Type	L. sativa var. crispa	L. sativa var. capitata8	L. sativa var. angustana11	L. saligna9	L. virosa10	
Gene number	41,375	38,919	40,341	42,908	—	
Gene length (bp)	154,935,371	118,296,912	148,482,760	107,734,927	148,482,760	
Average gene length (bp)	3744.66	3,228.98	3,756.58	2,521.94	3,756.58	
Exon length (bp)	73,199,166	67,064,559	48,098,291	55,886,651	48,098,291	
Average exon length (bp)	1769.16	1,830.56	1,216.88	1,308.24	1,216.88	
Exon number	208,636	204,681	238,700	213,805	238,700	
Average exon number	5.04	5.59	6.04	5	6.04	
CDS length (bp)	64,953,960	46,067,569	41,331,021	47,458,284	41,331,021	
Average CDS length (bp)	1569.88	1,257.44	1,045.67	1,110.94	1,045.67	
CDS number	202,518	172,388	231,105	209,929	231,105	
Average CDS number	4.89	4.71	5.85	4.91	5.85	
Intron length (bp)	81,736,205	51,232,353	100,384,469	51,848,276	100,384,469	
Average intron length (bp)	1975.5	1,398.42	2,539.71	1,213.71	2,539.71	
Intron number	167,261	168,045	199,174	171,086	199,174	
Average intron number	4.04	4.59	5.04	4	5.04	
Annotated gene number	41,004	31,348	34,966	40,730	39,887	
Annotated ratio	99.10%	80.55%	86.68%	94.9%	—	

Gene functions were inferred by aligning to the National Center for Biotechnology Information (NCBI) Non-Redundant (NR), EggNOG31, KOG, TrEMBL32, InterPro33 and Swiss-Prot32 protein databases using Diamond blastp (diamond v2.0.4.142) and the Kyoto Encyclopedia of Genes and Genomes (KEGG) database34 with an E-value threshold of 1E-5. Protein domains were annotated with InterProScan (v5.34-73.0)35, while motifs and domains within gene models were identified using PFAM databases36. Gene Ontology (GO) IDs for each gene were obtained from TrEMBL, InterPro and EggNOG. Approximately 41,004 (99.10%) of the predicted protein-coding genes in Green Elegance could be functionally annotated with known genes, conserved domains, and Gene Ontology terms (Table 6). This high annotation ratio (99.10%) is the highest among five Lactuca plants, including L. sativa var. capitata cv. Salinas, L. sativa var. angustana, L. saligna, and L. virosa (Table 6).

Whole genome synteny analysis

For synteny analysis, genomes of four other Lactuca species, including L. sativa var. capitata, L. sativa var. angustana, L. virosa and L. saligna assemblies, were aligned to the L. sativa var. crispa genome using Mummer (v 4.0)37 with the parameters: -c 500 -b 500 -l 100–maxmatch (Fig. 6). Raw alignment results were filtered using delta filter with parameters: -1 -i 90 -l 500. MCScanX identified syntenic blocks38 with the parameter -s 15 (number of genes required to call a collinear block) and visualized them using jcvi v1.2.839 with the parameter–minspan = 30.Fig. 6 Whole genome synteny among five genome assemblies in the genus Lactuca. Conserved syntenic blocks are highlighted with lines of different colors, corresponding to the nine pseudo-chromosomes.

Data Records

The raw genomic sequencing data used for genome assembly are available in the Genome Sequence Archive (GSA)40 in the National Genomics Data Center (NGDC), Beijing Institute of Genomics (China National Center for Bioinformation)41, Chinese Academy of Sciences (https://bigd.big.ac.cn/gsa). The accession number CRA01487342 covers genome survey data, transcriptomic sequencing data, PacBio HiFi sequencing data, ONT sequencing data, and Hi-C sequencing data. The genome assembly and annotation files are available in the Genome Warehouse (GWH)43 in NGDC (accession number is GWHERDY0000000044), Genebank (JBFTWI000000000)45 and Figshare (10.6084/m9.figshare.25116548)46.

Technical Validation

To evaluate the completeness of L. sativa var. crispa cv. Green Elegance (version 1.2) assembly, Illumina short-read and PacBio long-reads data were mapped back to the assembly. The alignment was analyzed using Qualimap v.2.2.2. The mapping rate for both libraries was 99.75% (an average 48× coverage) for Illumina short reads and 99.85% (average 42× coverage) for PacBio long-reads. BUSCO v5.2.247 with OrthoDB was used to assess genome completeness. In genome syntenic analysis, L. sativa var. crispa, L. sativa var. capitata and L. sativa var. angustana showed high conservation, compelling evidence that the gross genome structure has been accurately assembled (Fig. 6). We have observed relatively few genomic arrangement ambiguities in the Hi-C contact heat map, though with some discontinuities, which were probably caused by highly repetitive sequences. Visually inspection of the Hi-C map also revealed that some points of ambiguity appeared to be centromeres, likely due to sequence similarity in these regions. Meanwhile, clear antidiagonals for several chromosomes were also observed in the Hi-C contact heat map, such as chromosome 7 and chromosome 8. Such a pattern may suggest a Rabl configuration of the chromosomes, which could be validated in future cytological investigations. Overall, 98.39% BUSCOs were complete and 0.50% fragmented in the assembled genome (Table 7). CEGMA (Core Eukaryotic Genes Mapping Approach) (v2.5) analysis showed that 99.78% (457 CEG, Core Eukaryotic Genes) of CEGMA genes were present in the genome48. The LTR Assembly Index (LAI)49 of 17.34 indicated a high-quality genome assembly for L. sativa var. crispa cv. Green Elegance, with better continuity and completeness compared to other Lactuca species (Table 7). The higher ratio of complete BUSCOs and LAI values, compared to the other four Lactuca species, indicate the superior quality of the genome assembly for L. sativa var. crispa cv. Green Elegance (v1.2).Table 7 BUSCO and LAI assessments of L. sativa var. crispa, L. sativa var. capitata, L. sativa var. angustana, L. saligna, and L. virosa.

Type	L. sativa var. crispa	L. sativa var. capitata8	L. sativa var. angustana11	L. saligna9	L. virosa10	
Complete BUSCOs	1588 (98.39%)	1589 (98.45%)	1586 (98.27%)	1495 (92.63%)	1565 (96.96%)	
Complete and single-copy BUSCOs (S)	1551 (96.10%)	1554 (96.28%)	1535 (95.11%)	1460 (90.46%)	1507 (93.37%)	
Complete and duplicated BUSCOs (D)	37 (2.29%)	35 (2.17%)	51 (3.16%)	35 (2.17%)	58 (3.59%)	
Fragmented BUSCOs (F)	8 (0.50%)	3 (0.19%)	8 (0.50%)	14 (0.87%)	10 (0.62%)	
Missing BUSCOs (M)	18 (1.12%)	22 (1.36%)	20 (1.24%)	105 (6.51%)	39 (2.42%)	
Total Lineage BUSCOs	1,614	1,614	1,614	1,614	1,614	
LAI	17.34	18.13	—	—	—	

Acknowledgements

This research was supported by the Key Project at Central Government Level: The Ability Establishment of Sustainable Use for Valuable Chinese Medicine Resources (2060302), the Innovation and Development Program of Beijing Vegetable Research Center (KYCX202304), Beijing Joint Research Program for Germplasm Innovation and New Variety Breeding (G20220628003-01), and Collaborative Innovation Program of Beijing Vegetable Research Center (XTCX202302).

Author contributions

D.L., J.T., and B.Z. designed and coordinated the study; X.L., H.D., Y.Y., C.W., Z.X., and B.Z. collected and prepared plant samples; Y.X., J.T., and C.S. performed the bioinformatic analyses; B.Z. and D.L. drafted the manuscript; J.T., J.Z. and C.S. revised the manuscript. All authors approved the final manuscript.

Code availability

No custom code was used for this study. All data analyses were conducted using published bioinformatics software with default settings unless otherwise specified.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Wei T Whole-genome resequencing of 445 Lactuca accessions reveals the domestication history of cultivated lettuce Nat. Genet. 2021 53 752 760 10.1038/s41588-021-00831-0 33846635
Wei, T. et al. Whole-genome resequencing of 445 Lactuca accessions reveals the domestication history of cultivated lettuce. Nat. Genet. 53, 752–760, 10.1038/s41588-021-00831-0 (2021).33846635 10.1038/s41588-021-00831-0
2. Lindqvist K On the origin of cultivated lettuce Hereditas 1960 46 319 350 10.1111/j.1601-5223.1960.tb03091.x
Lindqvist, K. On the origin of cultivated lettuce. Hereditas 46, 319–350, 10.1111/j.1601-5223.1960.tb03091.x (1960).10.1111/j.1601-5223.1960.tb03091.x
3. de Vries IM Origin and domestication of Lactuca sativa L Genet. Resour. Crop Evol. 1997 44 165 174 10.1023/A:1008611200727
de Vries, I. M. Origin and domestication of Lactuca sativa L. Genet. Resour. Crop Evol. 44, 165–174, 10.1023/A:1008611200727 (1997).10.1023/A:1008611200727
4. Zohary D The wild genetic resources of cultivated lettuce (Lactuca sativa L.) Euphytica 1991 53 31 35 10.1007/BF00032029
Zohary, D. The wild genetic resources of cultivated lettuce (Lactuca sativa L.). Euphytica 53, 31–35, 10.1007/BF00032029 (1991).10.1007/BF00032029
5. Křístková E Doležalová I Lebeda A Vinter V Novotná A Description of morphological characters of lettuce (Lactuca sativa L.) genetic resources. A review Hortic. Sci.e 2018 35 113 129 10.17221/4/2008-HORTSCI
Křístková, E., Doležalová, I., Lebeda, A., Vinter, V. & Novotná, A. Description of morphological characters of lettuce (Lactuca sativa L.) genetic resources. A review. Hortic. Sci.e 35, 113–129 (2018).10.17221/4/2008-HORTSCI
6. Lebeda, A., Ryder, E. J., Sideman, R., Ivana, D. & Křístková, E.in Genetic resources, chromosome engineering, and crop improvement Vol. 3 (ed R. J. Singh) 377–472 (2006).
7. Zhang L RNA sequencing provides insights into the evolution of lettuce and the regulation of flavonoid biosynthesis Nat. Commun. 2017 8 2264 10.1038/s41467-017-02445-9 29273740
Zhang, L. et al. RNA sequencing provides insights into the evolution of lettuce and the regulation of flavonoid biosynthesis. Nat. Commun. 8, 2264, 10.1038/s41467-017-02445-9 (2017).29273740 10.1038/s41467-017-02445-9
8. Reyes-Chin-Wo S Genome assembly with in vitro proximity ligation data and whole-genome triplication in lettuce Nat. Commun. 2017 8 14953 10.1038/ncomms14953 28401891
Reyes-Chin-Wo, S. et al. Genome assembly with in vitro proximity ligation data and whole-genome triplication in lettuce. Nat. Commun. 8, 14953, 10.1038/ncomms14953 (2017).28401891 10.1038/ncomms14953
9. Xiong W The genome of Lactuca saligna, a wild relative of lettuce, provides insight into non-host resistance to the downy mildew Bremia lactucae Plant J. 2023 115 108 126 10.1111/tpj.16212 36987839
Xiong, W. et al. The genome of Lactuca saligna, a wild relative of lettuce, provides insight into non-host resistance to the downy mildew Bremia lactucae. Plant J. 115, 108–126, 10.1111/tpj.16212 (2023).36987839 10.1111/tpj.16212
10. Xiong W Genome assembly and analysis of Lactuca virosa: implications for lettuce breeding G3-GENES GENOM GENET 2023 13 jkad204 10.1093/g3journal/jkad204
Xiong, W. et al. Genome assembly and analysis of Lactuca virosa: implications for lettuce breeding. G3-GENES GENOM GENET 13, jkad204, 10.1093/g3journal/jkad204 (2023).10.1093/g3journal/jkad204
11. Shen F Comparative genomics reveals a unique nitrogen-carbon balance system in Asteraceae Nat. Commun. 2023 14 4334 10.1038/s41467-023-40002-9 37474573
Shen, F. et al. Comparative genomics reveals a unique nitrogen-carbon balance system in Asteraceae. Nat. Commun. 14, 4334, 10.1038/s41467-023-40002-9 (2023).37474573 10.1038/s41467-023-40002-9
12. Ballouz S Dobin A Gillis JA Is it time to change the reference genome? Genome Biol. 2019 20 159 10.1186/s13059-019-1774-4 31399121
Ballouz, S., Dobin, A. & Gillis, J. A. Is it time to change the reference genome? Genome Biol. 20, 159, 10.1186/s13059-019-1774-4 (2019).31399121 10.1186/s13059-019-1774-4
13. Sun Y Shang L Zhu Q-H Fan L Guo L Twenty years of plant genome sequencing: achievements and challenges Trends Plant Sci. 2022 27 391 401 10.1016/j.tplants.2021.10.006 34782248
Sun, Y., Shang, L., Zhu, Q.-H., Fan, L. & Guo, L. Twenty years of plant genome sequencing: achievements and challenges. Trends Plant Sci. 27, 391–401, 10.1016/j.tplants.2021.10.006 (2022).34782248 10.1016/j.tplants.2021.10.006
14. Abu Almakarem AS Heilman KL Conger HL Shtarkman YM Rogers SO Extraction of DNA from plant and fungus tissues in situ BMC Res. Notes 2012 5 266 10.1186/1756-0500-5-266 22672795
Abu Almakarem, A. S., Heilman, K. L., Conger, H. L., Shtarkman, Y. M. & Rogers, S. O. Extraction of DNA from plant and fungus tissues in situ. BMC Res. Notes 5, 266, 10.1186/1756-0500-5-266 (2012).22672795 10.1186/1756-0500-5-266
15. Chen S Zhou Y Chen Y Gu J fastp: an ultra-fast all-in-one FASTQ preprocessor Bioinformatics 2018 34 i884 i890 10.1093/bioinformatics/bty560 30423086
Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884–i890 (2018).30423086 10.1093/bioinformatics/bty560
16. Marçais G Kingsford C A fast, lock-free approach for efficient parallel counting of occurrences of k-mers Bioinformatics 2011 27 764 770 10.1093/bioinformatics/btr011 21217122
Marçais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 764–770, 10.1093/bioinformatics/btr011 (2011).21217122 10.1093/bioinformatics/btr011
17. Pertea M StringTie enables improved reconstruction of a transcriptome from RNA-seq reads Nat. Biotechnol. 2015 33 290 295 10.1038/nbt.3122 25690850
Pertea, M. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat. Biotechnol. 33, 290–295, 10.1038/nbt.3122 (2015).25690850 10.1038/nbt.3122
18. Cheng H Concepcion GT Feng X Zhang H Li H Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm Nat. Methods 2021 18 170 175 10.1038/s41592-020-01056-5 33526886
Cheng, H., Concepcion, G. T., Feng, X., Zhang, H. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods 18, 170–175, 10.1038/s41592-020-01056-5 (2021).33526886 10.1038/s41592-020-01056-5
19. Burton JN Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions Nat. Biotechnol. 2013 31 1119 1125 10.1038/nbt.2727 24185095
Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat. Biotechnol. 31, 1119–1125, 10.1038/nbt.2727 (2013).24185095 10.1038/nbt.2727
20. Wolff J Galaxy HiCExplorer 3: a web server for reproducible Hi-C, capture Hi-C and single-cell Hi-C data analysis, quality control and visualization Nucleic Acids Res. 2020 48 W177 W184 10.1093/nar/gkaa220 32301980
Wolff, J. et al. Galaxy HiCExplorer 3: a web server for reproducible Hi-C, capture Hi-C and single-cell Hi-C data analysis, quality control and visualization. Nucleic Acids Res. 48, W177–W184, 10.1093/nar/gkaa220 (2020).32301980 10.1093/nar/gkaa220
21. Flynn JM RepeatModeler2 for automated genomic discovery of transposable element families Proc. Natl. Acad. Sci. USA 2020 117 9451 9457 10.1073/pnas.1921046117 32300014
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. USA 117, 9451–9457, 10.1073/pnas.1921046117 (2020).32300014 10.1073/pnas.1921046117
22. Ellinghaus D Kurtz S Willhoeft U LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons BMC Bioinform. 2008 9 18 10.1186/1471-2105-9-18
Ellinghaus, D., Kurtz, S. & Willhoeft, U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinform. 9, 18, 10.1186/1471-2105-9-18 (2008).10.1186/1471-2105-9-18
23. Xu Z Wang H LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons Nucleic Acids Res. 2007 35 W265 W268 10.1093/nar/gkm286 17485477
Xu, Z. & Wang, H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 35, W265–W268, 10.1093/nar/gkm286 (2007).17485477 10.1093/nar/gkm286
24. Ou S Jiang N LTR_retriever: a highly accurate and sensitive program for identification of long terminal repeat retrotransposons Plant Physiol. 2018 176 1410 1422 10.1104/pp.17.01310 29233850
Ou, S. & Jiang, N. LTR_retriever: a highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 176, 1410–1422, 10.1104/pp.17.01310 (2018).29233850 10.1104/pp.17.01310
25. Tarailo-Graovac M Chen N Using RepeatMasker to identify repetitive elements in genomic sequences Curr. Protoc. Bioinformatics 2009 25 4.10.11 14.10.14 10.1002/0471250953.bi0410s25
Tarailo-Graovac, M. & Chen, N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinformatics 25, 4.10.11–14.10.14, 10.1002/0471250953.bi0410s25 (2009).10.1002/0471250953.bi0410s25
26. Benson G Tandem repeats finder: a program to analyze DNA sequences Nucleic Acids Res. 1999 27 573 580 10.1093/nar/27.2.573 9862982
Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 27, 573–580, 10.1093/nar/27.2.573 (1999).9862982 10.1093/nar/27.2.573
27. Beier S Thiel T Münch T Scholz U Mascher M MISA-web: a web server for microsatellite prediction Bioinformatics 2017 33 2583 2585 10.1093/bioinformatics/btx198 28398459
Beier, S., Thiel, T., Münch, T., Scholz, U. & Mascher, M. MISA-web: a web server for microsatellite prediction. Bioinformatics 33, 2583–2585, 10.1093/bioinformatics/btx198 (2017).28398459 10.1093/bioinformatics/btx198
28. Stanke M Steinkamp R Waack S Morgenstern B AUGUSTUS: a web server for gene finding in eukaryotes Nucleic Acids Res. 2004 32 W309 W312 10.1093/nar/gkh379 15215400
Stanke, M., Steinkamp, R., Waack, S. & Morgenstern, B. AUGUSTUS: a web server for gene finding in eukaryotes. Nucleic Acids Res. 32, W309–W312, 10.1093/nar/gkh379 (2004).15215400 10.1093/nar/gkh379
29. Kim D Langmead B Salzberg SL HISAT: a fast spliced aligner with low memory requirements Nat. Methods 2015 12 357 360 10.1038/nmeth.3317 25751142
Kim, D., Langmead, B. & Salzberg, S. L. HISAT: a fast spliced aligner with low memory requirements. Nat. Methods 12, 357–360, 10.1038/nmeth.3317 (2015).25751142 10.1038/nmeth.3317
30. Grabherr MG Trinity: reconstructing a full-length transcriptome without a genome from RNA-seq data Nat. Biotechnol. 2013 29 644 652 10.1038/nbt.1883
Grabherr, M. G. et al. Trinity: reconstructing a full-length transcriptome without a genome from RNA-seq data. Nat. Biotechnol. 29, 644–652, 10.1038/nbt.1883 (2013).10.1038/nbt.1883
31. Huerta-Cepas J eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses Nucleic Acids Res. 2019 47 D309 D314 10.1093/nar/gky1085 30418610
Huerta-Cepas, J. et al. eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res. 47, D309–D314, 10.1093/nar/gky1085 (2019).30418610 10.1093/nar/gky1085
32. Boeckmann B The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003 Nucleic Acids Res. 2003 31 1 365 370 10.1093/nar/gkg095 12520024
Boeckmann, B. et al. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003. Nucleic Acids Res. 31(1), 365–370 (2003).12520024 10.1093/nar/gkg095
33. Mitchell A The InterPro protein families database: the classification resource after 15 years Nucleic Acids Res. 2015 43 D213 D221 10.1093/nar/gku1243 25428371
Mitchell, A. et al. The InterPro protein families database: the classification resource after 15 years. Nucleic Acids Res. 43, D213–D221, 10.1093/nar/gku1243 (2015).25428371 10.1093/nar/gku1243
34. Kanehisa M Goto S Sato Y Furumichi M Tanabe M KEGG for integration and interpretation of large-scale molecular data sets Nucleic Acids Res. 2012 40 D109 D114 10.1093/nar/gkr988 22080510
Kanehisa, M., Goto, S., Sato, Y., Furumichi, M. & Tanabe, M. KEGG for integration and interpretation of large-scale molecular data sets. Nucleic Acids Res. 40, D109–D114, 10.1093/nar/gkr988 (2012).22080510 10.1093/nar/gkr988
35. Jones P InterProScan 5: genome-scale protein function classification Bioinformatics 2014 30 1236 1240 10.1093/bioinformatics/btu031 24451626
Jones, P. et al. InterProScan 5: genome-scale protein function classification. Bioinformatics 30, 1236–1240, 10.1093/bioinformatics/btu031 (2014).24451626 10.1093/bioinformatics/btu031
36. Finn RD Pfam: the protein families database Nucleic Acids Res. 2014 42 D222 D230 10.1093/nar/gkt1223 24288371
Finn, R. D. et al. Pfam: the protein families database. Nucleic Acids Res. 42, D222–D230, 10.1093/nar/gkt1223 (2014).24288371 10.1093/nar/gkt1223
37. Marçais G MUMmer4: a fast and versatile genome alignment system PLoS Comput. Biol. 2018 14 e1005944 10.1371/journal.pcbi.1005944 29373581
Marçais, G. et al. MUMmer4: a fast and versatile genome alignment system. PLoS Comput. Biol. 14, e1005944, 10.1371/journal.pcbi.1005944 (2018).29373581 10.1371/journal.pcbi.1005944
38. Wang Y MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity Nucleic Acids Res. 2012 40 e49 e49 10.1093/nar/gkr1293 22217600
Wang, Y. et al. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Res. 40, e49–e49, 10.1093/nar/gkr1293 (2012).22217600 10.1093/nar/gkr1293
39. Tang H Synteny and collinearity in plant genomes Science 2008 320 486 488 10.1126/science.1153917 18436778
Tang, H. et al. Synteny and collinearity in plant genomes. Science 320, 486–488, 10.1126/science.1153917 (2008).18436778 10.1126/science.1153917
40. Chen T The Genome sequence archive family: toward explosive data growth and diverse data types Genomics Proteom. Bioinform. 2021 19 578 583 10.1016/j.gpb.2021.08.001
Chen, T. et al. The Genome sequence archive family: toward explosive data growth and diverse data types. Genomics Proteom. Bioinform. 19, 578–583, 10.1016/j.gpb.2021.08.001 (2021).10.1016/j.gpb.2021.08.001
41. Members C-N Partners Database resources of the National Genomics Data Center, China National Center for Bioinformation in 2023 Nucleic Acids Res. 2023 51 D18 D28 10.1093/nar/gkac1073 36420893
Members, C.-N. & Partners Database resources of the National Genomics Data Center, China National Center for Bioinformation in 2023. Nucleic Acids Res. 51, D18–D28, 10.1093/nar/gkac1073 (2023).36420893 10.1093/nar/gkac1073
42. 2024 NGDC Genome Sequence Archive https://ngdc.cncb.ac.cn/gsa/browse/CRA014873
NGDC Genome Sequence Archive. https://ngdc.cncb.ac.cn/gsa/browse/CRA014873 (2024).
43. Chen M Genome warehouse: a public repository housing genome-scale data Genomics Proteom. Bioinform. 2021 19 584 589 10.1016/j.gpb.2021.04.001
Chen, M. et al. Genome warehouse: a public repository housing genome-scale data. Genomics Proteom. Bioinform. 19, 584–589, 10.1016/j.gpb.2021.04.001 (2021).10.1016/j.gpb.2021.04.001
44. 2024 NGDC Genome Warehouse https://ngdc.cncb.ac.cn/gwh/Assembly/83750/show
NGDC Genome Warehouse. https://ngdc.cncb.ac.cn/gwh/Assembly/83750/show (2024).
45. 2024 NCBI GenBank JBFTWI000000000
NCBI GenBank. https://identifiers.org/ncbi/insdc:JBFTWI000000000 (2024).
46. Zhang B 2024 Gemome assembly and gene annotation files of Lactuca sativa var. crispacrispa cv. Green Elegance figshare 10.6084/m9.figshare.25116548
Zhang, B. Gemome assembly and gene annotation files of Lactuca sativa var. crispa cv. Green Elegance. figshare.10.6084/m9.figshare.25116548 (2024).10.6084/m9.figshare.25116548
47. Simão FA Waterhouse RM Ioannidis P Kriventseva EV Zdobnov EM BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs Bioinformatics 2015 31 3210 3212 10.1093/bioinformatics/btv351 26059717
Simão, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210–3212, 10.1093/bioinformatics/btv351 (2015).26059717 10.1093/bioinformatics/btv351
48. Parra G Bradnam K Korf I CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes Bioinformatics 2007 23 1061 1067 10.1093/bioinformatics/btm071 17332020
Parra, G., Bradnam, K. & Korf, I. CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics 23, 1061–1067, 10.1093/bioinformatics/btm071 (2007).17332020 10.1093/bioinformatics/btm071
49. Ou S Chen J Jiang N Assessing genome assembly quality using the LTR Assembly Index (LAI) Nucleic Acids Res. 2018 46 e126 e126 10.1093/nar/gky730 30107434
Ou, S., Chen, J. & Jiang, N. Assessing genome assembly quality using the LTR Assembly Index (LAI). Nucleic Acids Res. 46, e126–e126, 10.1093/nar/gky730 (2018).30107434 10.1093/nar/gky730
