
==== Front
Plant Commun
Plant Commun
Plant Communications
2590-3462
Elsevier

S2590-3462(24)00291-8
10.1016/j.xplc.2024.100983
100983
Resource Article
Streamlined whole-genome genotyping through NGS-enhanced thermal asymmetric interlaced (TAIL)-PCR
Zhao Sheng 1
Wang Yue 1
Zhu Zhenghang 1
Chen Peng 1
Liu Wuge 2
Wang Chongrong 2
Lu Hong 1
Xiang Yong 1
Liu Yuwen 1
Qian Qian qianqian188@hotmail.com
1∗
Chang Yuxiao changyuxiao@caas.cn
1∗∗
1 Shenzhen Branch, Guangdong Laboratory of Lingnan Modern Agriculture, Genome Analysis Laboratory of the Ministry of Agriculture and Rural Affairs, Agricultural Genomics Institute at Shenzhen, Chinese Academy of Agricultural Sciences, Shenzhen 518120, China
2 Guangdong Provincial Key Laboratory of New Technology in Rice Breeding, Rice Research Institute, Guangdong Academy of Agricultural Sciences, Guangzhou 510640, China
∗ Corresponding author qianqian188@hotmail.com
∗∗ Corresponding author changyuxiao@caas.cn
05 6 2024
09 9 2024
05 6 2024
5 9 10098329 1 2024
21 4 2024
2 6 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Whole-genome genotyping (WGG) stands as a pivotal element in genomic-assisted plant breeding. Nevertheless, sequencing-based approaches for WGG continue to be costly, primarily owing to the high expenses associated with library preparation and the laborious protocol. During prior development of foreground and background integrated genotyping by sequencing (FBI-seq), we discovered that any sequence-specific primer (SP) inherently possesses the capability to amplify a massive array of stable and reproducible non-specific PCR products across the genome. Here, we further improved FBI-seq by replacing the adapter ligated by Tn5 transposase with an arbitrary degenerate (AD) primer. The protocol for the enhanced FBI-seq unexpectedly mirrors a simplified thermal asymmetric interlaced (TAIL)-PCR, a technique that is widely used for isolation of flanking sequences. However, the improved TAIL-PCR maximizes the primer-template mismatched annealing capabilities of both SP and AD primers. In addition, leveraging of next-generation sequencing enhances the ability of this technique to assay tens of thousands of genome-wide loci for any species. This cost-effective, user-friendly, and powerful WGG tool, which we have named TAIL-PCR by sequencing (TAIL-peq), holds great potential for widespread application in breeding programs, thereby facilitating genome-assisted crop improvement.

This study presents a simple, affordable, sequencing-based method for genome-wide genotyping based on single TAIL-PCR that uses both sequence-specific and arbitrary degenerate primers. This new genotyping method can be applied to a wide variety of molecular breeding projects.

.

Key words

whole-genome genotyping
primer-template mismatched annealing
specific primer
arbitrary degenerate primer
molecular breeding
Published: June 5, 2024
==== Body
pmcIntroduction

Rapid iterations in third-generation sequencing technology have facilitated the assembly and release of over 800 plant reference genomes to the public (Marks et al., 2021; Sun et al., 2022). The availability of high-quality genome sequences (Zhang et al., 2022, 2024; Lu et al., 2024b), sometimes even at the telomere-to-telomere level (Shang et al., 2023; Sun et al., 2023; Yang et al., 2023), has empowered plant breeders to transition from conventional ﬁeld selection processes to genome-assisted breeding (Xu et al., 2022).

In addition to high-quality reference genomes, modern plant breeding relies on whole-genome genotyping (WGG) of thousands to hundreds of thousands of crop samples (Davey et al., 2011). Five methods are commonly used for WGG: whole-genome sequencing (WGS), DNA microarrays (Rasheed et al., 2017), hybrid capture sequencing (Hodges et al., 2009; Guo et al., 2021), restriction-site-associated DNA sequencing (RADseq) (Baird et al., 2008; Andrews et al., 2016), and multiple PCR (Markoulatos et al., 2002), with its derivatives like foreground and background integrated genotyping by sequencing (FBI-seq) (Zhao et al., 2023), Hyper-seq (Zou and Xia, 2022), and multiplex restriction amplicon sequencing (MRASeq; Bernardo et al., 2020). Each of these methods has limitations that hinder widespread application, especially for minor crops. WGS is the most robust and scalable approach. However, for genetic research or breeding samples, sequencing only ∼1%–10% of the genome is sufficient; WGS thus incurs unnecessary costs in the sequencing step. The other four methods inherently belong to the same category of reduced representation genome assay, aimed at selecting a portion of the genome (typically 1%–10%) to represent the whole-genome information. The difference lies in the selection methods. DNA microarrays and hybrid capture sequencing use probe hybridization in either solid (Rasheed et al., 2017) or solution phases (Kozarewa et al., 2015), whereas multiple PCR selects genomic loci through PCR amplification (Markoulatos et al., 2002). All require an initial investment for the development and optimization of target probes or primers. These approaches are particularly well suited for high-demand crops, as a large sample cohort can effectively reduce the marginal cost per sample for oligo development and optimization. RADseq involves selection of the fragment-associated restriction site and does not require initial investment for probes; however, it involves numerous experimental steps and is challenging to automate, minimizing its utility for high-throughput WGG (Andrews et al., 2016). Therefore, we still lack a cost-effective and robust WGG method for breeding projects.

Next-generation sequencing (NGS) prices have declined ∼100 000 fold since the initial introduction of this technology in 2008 (http://www.genome.gov/sequencingcostsdata), but DNA library preparation remains an expensive prerequisite for NGS. Therefore, the key challenge for WGG within a breeding project is the robust and simplified selection of genome fragments, coupled with library preparation for addition of sequencing adapters.

We previously demonstrated that in addition to annealing to its specific locus, a specific primer (SP) could also anneal to genomic DNA (gDNA) templates, initiating genome-wide PCR amplification at tens of thousands of loci through primer-template mismatched annealing (PTMA; Zhao et al., 2023). This unexpected capability of SPs was reminiscent of thermal asymmetric interlaced (TAIL)-PCR (Liu and Whittier, 1995; Liu and Chen, 2007), in which an SP combined with an arbitrary degenerate (AD) primer is used to amplify the unknown flanking sequence of the SP locus. We noted that the PCR product of primary TAIL-PCR frequently appears as a smear (Liu and Whittier, 1995; Liu and Chen, 2007), resembling an NGS library. The smear in the primary TAIL-PCR product, along with the unexpected ability of SPs for PTMA at tens of thousands of loci, prompted us to reconsider the composition of the smear. If the SP primer can also initiate genome-wide PCR amplification at tens of thousands of loci in TAIL-PCR, similar to FBI-seq (Zhao et al., 2023), then perhaps the smear in the primary TAIL-PCR product originates from amplification across tens of thousands of loci. This aligns with the objectives of multiple PCR or hybrid capture sequencing for fragment selection. However, the fragment selection and enrichment by TAIL-PCR appeared to be simpler compared with multiple PCR (Markoulatos et al., 2002) and hybrid capture sequencing (Hodges et al., 2009; Guo et al., 2021), making it perhaps the most straightforward method. In the present study, we demonstrate that WGG through TAIL-PCR by sequencing (TAIL-peq) is feasible. Our method enables the amplification of several hundred thousand loci across the genome of any species in a single PCR reaction, necessitating only an investment in SP and AD primers. Consequently, our study presents a novel, efficient approach for high-throughput WGG of numerous species.

Results

TAIL-peq workflow and amplicon characteristics

The workflow of TAIL-peq comprises two steps (Figure 1). First, we perform TAIL-PCR using barcoded SP and AD primers (Figure 1A). After PCR, fragments selected from each sample are assigned specific barcodes. Second, these barcoded fragments are pooled together for the next phase of dual-index TruSeq library preparation and sequencing (Figure 1B).Figure 1 TAIL-PCR by sequencing (TAIL-peq) workflow.

(A) In primary TAIL-PCR amplification, SP and AD primer pairs ligated with distinct barcodes bind with the gDNA via PTMA. After amplification of each sample, two specific barcodes are attached at both ends of the amplicons.

(B) Barcoded amplicons from all samples are pooled into one tube and then subjected to typical TruSeq library preparation steps, including end-repair, A-tailing, Y-shaped adapter ligation, library amplification, and sequencing.

To clearly demonstrate the amplification that took place in TAIL-PCR using rice Nipponbare gDNA as a template, we performed three technical replicates with three SPs (designed for the rice genes ALK, DTH2, Pi2 and containing the same F01barcode; Supplemental Table 1), combined with the same AD primer (AD1_2304; here, 2304 refers to the degeneracy level). The combinations of AD1_2304 with ALK_F01, DTH2_F01, and Pi2_F01 were referred to as primer pairs A, B, and C, respectively (Supplemental Table 2). To avoid missing details of the TAIL-PCR amplification result, we prepared the TruSeq library using the total PCR product of each replicate. After sequencing, we obtained 1.7–14.3 Gb of data per replicate, from which 1.0 Gb of data was randomly subsampled for further analyses.

We found five categories of amplicons, distinguished by their initiating primers at both ends. Amplicons produced by the SP + AD combination accounted for the largest proportion of reads (Figure 2A; Supplemental Table 2), ranging from 90.2% to 90.8% for primer pair A, 67.1% to 68.1% for pair B, and 79.6% to 81.3% for pair C. The remaining amplicons included amplifications derived from SP + SP, AD + AD, and single SP or AD primers.Figure 2 Evaluation of the TAIL-peq genotyping method.

Three technical replicates were performed for each primer pair, A–C. After sequencing, 1.0 Gb of data were randomly subsampled from each replicate of primer pairs A–C for further analysis. Primer pairs A, B, and C correspond to ALK_F01 + AD1_2304, DTH2_F01 + AD1_2304, and Pi2_F01 + AD1_2304, respectively.

(A) Proportion of reads generated from five amplicon categories (SP + AD, SP + SP, AD + AD, and SP or AD alone).

(B) The distribution of TAIL-peq reads across the rice genome revealed a pattern in which reads accumulated, forming distinct tags.

(C) Average sequencing depth of TAIL-peq tags generated with primer pairs A–C.

(D) Number of shared TAIL-peq tags among three technical replicates with primer pair A.

(E) Heatmap displaying the distribution of 11 739 shared tags (from three technical replicates of primer pair A) across the rice genome. Distributions were calculated within 0.1-Mb windows.

(F) Evaluation of the breadth of coverage, defined as the percentage of bases that were sequenced a given number of times (here 1×–51×) in each of three technical replicates with primer pair A.

(G) Lorenz curves of three technical replicates generated with primer pair A. The Lorenz curve illustrates the cumulative fraction of reads versus the cumulative fraction of the genome. The black dotted line represents theoretical perfectly uniform coverage. TAIL-peq data from the three technical replicates exhibited similar deviations from the diagonal, demonstrating the replicability of this method.

(H) High SNP accuracy calculated at various coverage depth thresholds (1×–10×) in each of three technical replicates with primer pair A.

The sequencing data were subsequently mapped to the Nipponbare_v7.0 reference genome, and the reads tended to pile up (multiple reads mapped to the same region), forming genome-wide sequence tags (Figure 2B; Supplemental Figures 1 and 2). As the amount of sequencing data increased from 0.1 to 12.0 Gb, we observed a gradual increase in genome coverage, which remained below 5% (Supplemental Figure 3). By contrast, WGS data within the same range achieved genome coverage ranging from 19.6% to 96.3% (Supplemental Figure 3A). TAIL-peq thus amplified only a portion of the genome and achieved the objective of reducing genome representation. With 1.0 Gb of data, we identified 18 036–18 376, 105 429–109 247, and 66 091–69 421 tags with primer pairs A–C, respectively (Supplemental Table 2). The average sequencing depths of these tags were 264.6–268.2×, 44.8–45.3×, and 72.1–76.4× (Figure 2C). These differences aligned with the variation in genome coverage, suggesting primer-pair-specific disparities in PCR amplification capacity. Specifically, primer pair A had the fewest annealing sites, resulting in the highest read depth but the lowest tag number. By contrast, pair B had the most annealing sites, leading to the lowest read depth but the highest tag number. Overall, these findings indicate that TAIL-peq selectively enriches specific genomic fragments, replicating the effects achieved through hybrid capture sequencing or multiple PCR amplification.

TAIL-peq replicability and accuracy

In contrast to methods like multiple PCR, DNA array, and hybrid capture sequencing that use specific probes/primers to select target DNA fragments, TAIL-peq uses PCR amplification primed by SP and AD primers. Consequently, it is essential to assess its reproducibility. We tested this by comparing the number of sequence tags shared across three technical replicates for each primer pair. With 1.0 Gb of sequencing data, pairs A–C exhibited 11 739 (63.9%–65.1%), 54 585 (50.0%–51.8%), and 35 176 (50.7%–53.2%) shared tags (defined as a missing rate of 0%), respectively (Figure 2D; Supplemental Figure 4; Supplemental Table 2). These shared tags were distributed across the rice genome (Figure 2E; Supplemental Figure 5). We then evaluated the breadth of coverage, defined as the percentage of bases that were sequenced a given number of times (Sims et al., 2014). For each of the three primer pairs, the breadth of coverage at the same depth was highly consistent among replicates (Figure 2F; Supplemental Figure 6). Using a sliding-window approach, we found that the correlation coefficients for the mean sequencing depth for each primer pair exceeded 97.2% (Supplemental Table 2). Finally, we plotted Lorenz curves to represent the cumulative fraction of total reads versus the cumulative fraction of the genome and observed high reproducibility across the three replicates for each primer pair, with only minor deviations (Figure 2G; Supplemental Figure 7). Overall, these findings underscore the robust replicability of TAIL-peq data.

Upon alignment of 1.0 Gb of sequencing data to the R498 reference genome (Du et al., 2017), we detected 11 633 ± 305, 54 887 ± 2667, and 35 315 ± 1014 SNPs across the three replicates of each primer pair (A–C). We observed that SNP accuracy, validated by whole-genome re-sequencing, progressively increased with the SNP depth threshold. Specifically, the accuracy reached 98.56% ± 0.26%, 99.06% ± 0.20%, and 99.39% ± 0.21% for pairs A–C, respectively, even at a minimal depth of 1× (Figure 2H; Supplemental Figure 8). These results demonstrate the high accuracy of the TAIL-peq genotyping method.

Effects of SP and AD primers on TAIL-peq amplification

We investigated the effects of SP and AD primers on the amplicons generated with TAIL-peq. Six SPs were paired with two sets of AD primers (AD1 and AD2), each set with a similar sequence but three levels of degeneracy (2304, 768, and 144). Three technical replicates were generated for each primer pair, and 0.5 Gb of raw sequencing reads were randomly subsampled from each technical replicate for further analysis.

We first calculated the proportions of reads derived from the five categories of amplicons. Consistent with earlier results, different primer pairs produced reads that accounted for various proportions of the total. Of the six SPs analyzed, Rf3_F01-containing primer pairs produced the most consistent proportions of each amplicon category, followed by Rf4_F01, Rim2_F01, and Waxy_F01, whereas BADH2_F01- and Hd1_F01-containing primer pairs displayed the greatest variation (Figure 3A; Supplemental Table 3). Reads derived from SP + AD, AD + AD, or single AD primer extension typically comprised the largest proportion of reads (58.2%–93.1%), whereas SP + SP reads generally accounted for the lowest proportion (0.01%–7.6%). Importantly, using the unique barcodes of the SP and AD primers, more than 93.3% of the raw sequencing reads were resolvable for each primer pair (Figure 3A; Supplemental Table 3).Figure 3 The effect of SP and AD primer pairs on TAIL-peq amplification.

Three technical replicates were executed for each primer pair, and after sequencing, 0.5 Gb of data were randomly selected from each replicate for further analysis.

(A) Proportion of reads amplified from five amplicon categories.

(B) Total number of TAIL-peq tags detected with each SP and AD primer pair.

(C) Number of shared tags detected among three technical replicates.

(D) Heatmap displaying the distribution of 9782 shared tags (coverage depth threshold ≥30×) across the rice genome. The primer pair used was Hd1_F01 + AD1_2304. Tag numbers were calculated within 0.1-Mb windows.

We next calculated the number of tags detected using six SPs paired with each degeneracy level of AD1 and AD2 primers. Regardless of AD primer or degeneracy level, Waxy_F01-containing primer pairs still generated the largest number of tags (68 906–150 364), followed by BADH2_F01 (63 366–137 783), Hd1_F01 (37 708–84 367), Rf4_F01 (48 384–86 435), Rim2_F01 (28 032–66 899), and Rf3_F01 (22 113–53 298) (Figure 3B; Supplemental Table 3). More significant variations were observed in the tag counts generated by individual AD primers paired with different SP primers compared with the relatively consistent tag counts produced by individual SP primers across various AD primer combinations (Figure 3B). Within the three replicates of each primer pair, Waxy_F01- or BADH2_F01-containing pairs consistently had the largest number of shared tags (38 155–79 895), whereas Rf3_F01-containing primer pairs had the fewest (13 153–26 366) (Figure 3C; Supplemental Table 3). When the read depth threshold increased to 30× or 50×, thousands of shared tags were still detected for each primer pair (Supplemental Figure 9) and were distributed across all chromosomes of the rice genome (Figure 3D; Supplemental Figure 10). Remarkably, all Hd1_F01-containing primer pairs showed more than 2610 shared tags, even at a depth of 100× (Supplemental Figure 9; Supplemental Table 3), indicating that Hd1_F01 had the strongest binding stability among the 6 SPs. Together, these findings clearly revealed that both SP and AD primers had variable capacities for PTMA amplification and that the SP had a greater effect than the AD primer on the number of sequence tags produced. Nevertheless, the quantity of tags produced by each primer combination consistently stabilized within a specific order of magnitude, reaffirming its robustness and reproducibility.

Application of TAIL-peq to maize germplasm and rice genetic populations

We evaluated the efficacy of TAIL-peq for maize breeding and rice genetics studies. With respect to maize, we used a panel of 127 inbred maize lines, aiming to elucidate their inherent genetic connections. We used the primer sets Hd1_Bxx and AD2_Bxx, with Bxx indicating the barcode ID (Supplemental Table 1). The TAIL-PCR products were pooled in groups of 24 lines for subsequent TruSeq library preparation. After sequencing, we acquired an average of ∼2.2 ± 1.7 Gb of data per line, corresponding to an average genome coverage of 1.1% ± 0.5% (Figure 4A; Supplemental Table 4). From this dataset, we were able to pinpoint 4609 high-quality SNPs with missing rates less than 20% dispersed across the 10 chromosomes (Figure 4B). A neighbor-joining (NJ) phylogenetic tree constructed from the gathered SNPs demonstrated that the inbred lines could be categorized into three distinct clusters. These clusters corresponded to several prominent heterotic groups that have been used extensively in maize breeding endeavors in recent decades: Stiff Stalk (SS; 30 lines), Sweet (24 lines), and Iodent (66 lines) (Figure 4C). The lines in the Sweet and Iodent groups appeared to share a closer genetic relationship among themselves than those in the SS group. A principal-component analysis (PCA) of the same population sample revealed a consistent distribution among the three groups (Figure 4D). These findings, which are consistent with earlier germplasm evaluation results (Liu et al., 2003; Hu et al., 2021), suggest that TAIL-peq can be effectively used for studies of crop germplasm materials.Figure 4 Application of TAIL-peq in maize germplasm characterization.

(A) Genome coverage of the 127 maize inbred lines.

(B) Genome-wide distribution of the 4609 high-quality SNPs.

(C) Neighbor-joining tree of the 127 inbred lines based on high-quality SNPs.

(D) PCA depicting the genetic diversity within the maize population.

To evaluate the practical application of TAIL-peq to another crucial crop species, we extended our research by developing a linkage map for a rice genetic population of 220 recombinant inbred lines (RILs). On average, each sample produced ∼0.8 ± 0.9 Gb of sequencing data, equivalent to 1.2% ± 0.7% genome coverage (Figure 5A). Using these TAIL-peq data, we detected roughly 23 997.3 ± 15 239.8 SNPs for each line (Supplemental Table 5). The numerous SNPs paved the way for identification of recombination breakpoints in the population, which totaled 6471 and averaged 29.4 per line (Figure 5B). These data enabled the creation of a high-density genetic map comprising 2618 bin markers (Figure 5C; Supplemental Table 6), spanning ∼1302.4 cM with an average interval of 0.54 cM between markers. There was strong collinearity between the genetic and physical maps (Pearson’s r = 0.98) (Figure 5D; Supplemental Table 6). Collectively, these analyses of two important plant species attest to both the high quality of TAIL-peq data and its appropriateness for plant breeding and genetic research.Figure 5 Application of TAIL-peq to rice genetics research.

(A) Genome coverage assessment of the 220 rice RILs.

(B) Construction of a recombination bin map illustrating recombination events in the RIL population.

(C) Development of a high-density genetic map based on the population.

(D) Collinearity analysis between the genetic map and the physical map of Oryza sativa ssp. japonica cv. ‘Nipponbare’ based on the V7.0 reference genome. Bin markers positioned on linkage groups (LGs) are connected to their corresponding chromosome (Chr) positions with lines of various colors.

TAIL-peq is robust for whole-genome genotyping of any species

Because the amplification in TAIL-peq is initiated from PTMA of both SP and AD primers, genomes of all species theoretically possess a large number of potential annealing sites. We therefore hypothesized that SP and AD primer pairs could be used for WGG in any species, not merely the species for which the SPs were specifically designed. To verify this hypothesis, we deployed TAIL-peq across 10 distinct plant species: eggplant, green bean, lettuce, pepper, radish, sunflower, tomato, sorghum, soybean, and hexaploid wheat. These plants have genome sizes ranging from ∼0.46 Gb (radish) to ∼17 Gb (wheat). We carried out TAIL-peq library preparation with three technical replicates per species, generating between 0.52 and 2.2 Gb of sequencing data per replicate. Alignments to the respective reference genomes showed that the genome coverages varied from 0.5% to 0.6% in green bean to 6.8% to 7.3% in sorghum (Supplemental Table 7). Using these data, we identified ∼16.1–238.3K tags per sample, approximately half of which were shared among the three technical replicates of each species (Supplemental Table 7). Moreover, these shared tags were distributed across the whole genome of each species (Figure 6A and 6B; Supplemental Figure 11), thus demonstrating the robustness of TAIL-peq for WGG across various species.Figure 6 Genomic distribution of shared PTMA tags detected in two plant and two animal species.

(A–D) Tags were included in this analysis if they were shared among three technical replicates of a given sample and primer pair. There were a total of 84 275, 47 654, 99 190, and 101 285 shared tags in sunflower (A), wheat (B), dog (C), and pig (D). In (A) and (B), the primer pair was DTH2_F01 + AD1_2304; in (C) and (D), the primer pair was Hd1_F01 + AD2_2304. Tag numbers were calculated within 0.1-Mb windows.

Finally, we evaluated the efficacy of TAIL-peq in dog and pig. The animal subjects included a golden retriever (D_139), a corgi (D_231), and two Duroc pigs (P_421 and P_482). For each of the corresponding samples, the rice-derived SP (Hd1_F01) was paired with the AD2 primer backbone at three different degeneracy levels (2304, 768, and 144). By analyzing 1.0 Gb of randomly subsampled sequencing data from the total dataset, we identified approximately 100 000 tags per sample for both dog and pig (Supplemental Table 8). Notably, we observed that a reduction in AD2 degeneracy from 2304 to 144 led to a reduction of ∼30% in both genome coverage and tag number (Supplemental Table 8). Despite this reduction in genome coverage, the distribution of shared tags across the genomes of both dog and pig remained satisfactorily comprehensive (Figure 6C and 6D; Supplemental Figure 12–15). Overall, these findings suggest that TAIL-peq is a versatile method suitable for WGG in both plant and animal species.

Discussion

Over the past two decades, a series of advances in NGS technology have brought about a revolution in plant breeding research, opening up the genomics era of crop improvement (Shen et al., 2022). To date, chromosome-level reference genomes have been established for most major and minor crop species, and a number of WGG methods have been developed to fingerprint plants, greatly facilitating genome-assisted breeding. Despite advances in reducing sequencing costs and labor requirements, the expenses associated with library preparation continue to pose a barrier for many research groups, limiting their ability to perform large-scale genotyping for crop breeding. Here, we developed a novel WGG method, TAIL-peq, to address this issue. Our approach enables the amplification of tens of thousands of loci through stable PTMA in a single PCR. Barcoded amplicons from many samples can then be pooled to form just one sample for subsequent TruSeq library preparation steps. The entire TAIL-peq library preparation process is simple, requiring only basic infrastructure and molecular biology equipment and just one PCR reaction, making it one of the simplest WGG library preparation methods developed to date. Importantly, the overall cost of TAIL-peq library preparation for a single sample is approximately equivalent to that of a single TAIL-PCR (∼$1.75/sample; Supplemental Table 10).

On the one hand, our TAIL-peq approach can be considered an NGS-enhanced TAIL-PCR, given the unexpected revelation that the SP primer possesses an inherent capability to initiate PCR amplification across tens of thousands of loci in adapter-ligation PCR, facilitating genome-wide amplification (Zhao et al., 2023). On the other hand, the TAIL-PCR can also be regarded as an improvement of FBI-seq, in which the primer annealed to the adapter ligated by Tn5 transposase is replaced by an AD primer. This improvement avoids the use of Tn5 transposase and simplifies WGG library preparation to a single PCR amplification. In terms of SP and AD primer design, although both have demonstrated considerable sequence flexibility, we do not recommend the use of AD primers with a series of NNNNNN bases within the primer sequences. Instead, in the degenerate nucleotide positions of the AD primers, two or three should be designed to have three rather than four nucleotides to avoid self-pairing between the 3′ ends (Liu and Chen, 2007). In addition, we advise the use of a uniform annealing temperature across a genotyping project to minimize possible variations in PTMA amplification among different samples.

In terms of workflow complexity, monetary cost, and labor requirements, TAIL-peq presents distinct advantages compared with other WGG methods, including WGS, RADseq (especially the commonly used genotype by sequencing [Elshire et al., 2011; Wallace and Mitchell, 2017]), DNA chips, hybrid capture sequencing (Hodges et al., 2009; Guo et al., 2021), multiplex PCR (Markoulatos et al., 2002), FBI-seq (Zhao et al., 2023), Hyper-seq (Zou and Xia, 2022), and MRASeq (Bernardo et al., 2020) (Supplemental Tables 9 and 10). First, unlike WGS, RADseq, DNA chips, or FBI-seq, there is no need to fragment gDNA, which always involves a complex instrument (sonicator) or expensive reagents (Tn5 transposase or restriction endonucleases). Second, TAIL-peq does not require the laborious process of primer design, because PTMA allows direct use of the rice-specific primers in diverse species. Third, the pooled PCR-based library preparation approach is extremely efficient, enabling the preparation of a large number of samples in a short time within a limited budget. The convenience and cost-effectiveness of TAIL-peq for WGG library preparation, coupled with continuously decreasing sequencing costs, bring us closer to application in molecular breeding, particularly for minor crops.

In breeding programs, hundreds to thousands of well-distributed DNA markers are typically adequate for quantitative trait loci (QTL) mapping (Bernardo et al., 2020), germplasm resource evaluation, and kinship inference (Nguyen et al., 2020), and thousands to tens of thousands of molecular markers are used for genomic selection (Bhat et al., 2016). Here, we demonstrated that thousands to tens of thousands of genome-wide SNPs can be obtained by TAIL-peq genotyping, regardless of the primer pair tested or the species assayed. Moreover, TAIL-peq may be suitable for genome-wide association studies, which often require hundreds of thousands (or even more) of polymorphic molecular markers to achieve enhanced resolution and accuracy (Zhao et al., 2022; Lu et al., 2024a) following the imputation of missing genotype data, similar to how previous low-coverage sequencing data have been handled (Huang et al., 2012; Munyengwa et al., 2021; Watowich et al., 2023). Bi-parental genetic populations and natural populations of diverse accessions or inbred lines are extensively used in crop genetic improvement (Wurschum, 2012). TAIL-peq was here confirmed to enable SNP identification within a maize germplasm population, capturing variation that produced biologically relevant sample clusters in both PCA and NJ tree analyses. Furthermore, use of TAIL-peq in a rice RIL population produced a highly accurate linkage map of comparable quality to previous maps generated by WGS (Huang et al., 2009; Xie et al., 2010). In summary, we have developed an affordable, simple, powerful WGG method. This approach serves as a viable alternative to WGG for both plant and animal breeding programs. Our method is expected to enable large-scale WGG for plant breeding groups, particularly those with limited resources, ultimately contributing to targeted crop improvement.

Materials and methods

Biological materials and gDNA isolation

Seeds were obtained from the Agricultural Genomic Institute at Shenzhen, Chinese Academy of Agricultural Sciences (AGIS, CAAS), for the following plant materials: rice (Oryza sativa ssp. japonica cv. ‘Nipponbare’), eggplant (Solanum melongena), green bean (Phaseolus vulgaris), lettuce (Lactuca sativa), pepper (Capsicum annuum), radish (Raphanus sativus), sorghum (Sorghum bicolor), soybean (Glycine max ‘Tianlong1’), sunflower (Helianthus annuus), tomato (Solanum lycopersicum), and wheat (Triticum aestivum ‘Chinese Spring’). A rice RIL population developed from an N22 × Nipponbare cross was obtained from Prof. Xiang (AGIS, CAAS). A panel of maize inbred lines was randomly selected from a classical maize heterotic group obtained from Prof. Lu (AGIS, CAAS). Blood samples from two dogs (Canis lupus familiaris) and two pigs (Sus scrofa) were provided by Prof. Liu (AGIS, CAAS). Plant gDNA was extracted from seedling leaves using a modified cetyltrimethylammonium bromide protocol (Saghai-Maroof et al., 1984). Animal gDNA was extracted from the blood samples using the FastPure Blood DNA Isolation Mini Kit V2 (catalog no. DC111-01, Vazyme Biotech). The DNA quality and concentration were assessed by 0.8% agarose gel electrophoresis and a Qubit 4.0 fluorometer (Invitrogen), respectively.

Primer design

The SPs used for primary PCR amplification were designed for the rice genes ALK (Gao et al., 2003), DTH2 (Wu et al., 2013), Pi2 (Zhou et al., 2006), BADH2 (Bradbury et al., 2005), Hd1 (Nemoto et al., 2016), Rf3 (Zhang et al., 1997), Rf4 (Yao et al., 1997), Rim2 (Kwon et al., 2006), and Waxy (Cai et al., 1998) using Primer 5.0. The primers consisted of two sections: a short, unique barcode at the 5′ end and a gene-specific sequence at the 3′ end. AD primer sequences were designed on the basis of previously published primers (Liu and Chen, 2007) with minor modifications, namely the replacement of several bases to introduce three levels of degeneracy and the addition of unique barcodes at the 5′ ends. The barcodes of the primers in each SP and AD primer pair differed and were used to differentiate the samples. All primer sequences used in this study are listed in Supplemental Table 1.

Library preparation and sequencing

The primary TAIL-PCR reaction (Figure 1) was carried out in 50-μL mixtures, each containing 50 ng gDNA, 6 μL SP primer (10 μM), 20 μL AD primer (10 μM), 10 μL 5× TAB buffer, and 1 μL TruePrep Amplify Enzyme (catalog no. TD601-01, Vazyme Biotech). The thermocycling program was as follows: pre-denaturation at 98°C for 2 min; 5 cycles of 98°C for 30 s, 56°C for 30 s, and 72°C for 3 min; 7 cycles of 98°C for 30 s, 56°C for 30 s, 72°C for 3 min, 98°C for 30 s, 56°C for 30 s, 72°C for 3 min, 98°C for 30 s, 45°C for 30 s, and 72°C for 3 min; and a final extension at 72°C for 10 min. After primary PCR, products from multiple samples were pooled in a single tube using the all-in-one-seq method (Zhao et al., 2020). End-repair, A-tailing, short universal adapter ligation, and dual-indexing (index5 and index7) were carried out using the Hieff NGS 384 CDI Primer for Illumina (catalog no. 12412ES02, Yeasen Biotech) and the Hieff NGS Ultima Pro DNA Library Prep Kit for Illumina (catalog no. 12201ES96, Yeasen Biotech), following the manufacturer’s instructions. DNA library fragments of 400–800 bp were selected using an automatic cassette-based Pippin HT system (Sage Science). Sequencing was performed by Berry Genomics (Beijing, China) on the Illumina NovaSeq 6000 platform, producing paired-end 150-bp (PE150) reads.

For comparison with TAIL-peq sequencing results, a standard WGS library of rice Nipponbare gDNA was constructed with the TruePrep DNA Library Prep Kit V2 for Illumina (catalog no. TD501-01, Vazyme Biotech) following the manufacturer’s instructions. Size selection and sequencing were performed as described above.

Read de-multiplexing and tag identification

Raw sequencing reads were initially demultiplexed using the index5 and index7 sequences. Random subsampling of sequencing data was achieved using the “seqtk sample” command (v1.3-r106; https://github.com/lh3/seqtk). To determine the proportion of amplicons derived from each primer combination (SP + AD, SP + SP, AD + AD, SP alone, and AD alone) in each sample, the raw reads were divided with fastq-multx version 1.4.3 (https://github.com/brwnj/fastq-multx) using the barcode sequences located at the beginning of read1 and read2. Reads that did not contain barcode sequences were classified as unknown and discarded.

Barcode sequences were trimmed off with fastp version 0.23.0 (Chen et al., 2018), and reads were then aligned to the appropriate reference genome using BWA–MEM version 0.7.17 (Li and Durbin, 2009). Only reads with high mapping quality (MQ ≥ 30) were retained. SAMtools version 1.9 (Li et al., 2009) was used to calculate the sequencing depth at each base. Consistent with a previous study (Zhao et al., 2023), genomic loci were designated as sequence tags (∼120–150 bp) if they met the following criteria: (1) there were at least three reads beginning at the same coordinate; and (2) the sequencing depth at the beginning coordinate was at least 3× higher compared with the depth at the previous base. Tag distribution across the genome was visualized with the R package CMplot version 4.2.0 (https://github.com/YinLiLin/CMplot).

SNP calling and map construction

SNPs were called for the maize natural population sequenced with TAIL-peq. Reads with barcodes removed were aligned to the B73 reference genome (Jiao et al., 2017) using BWA–MEM version 0.7.17 (Li and Durbin, 2009). SAMtools version 1.9 (Li et al., 2009) was then used to convert SAM files to sorted BAM files with the following options: ‘view -bS -q 20.’ SNPs were detected using Genome Analysis Toolkit version 4.2.5.0 (McKenna et al., 2010). Genomic variants for each sample were identified in genomic variant call format (GVCF) using the HaplotypeCaller module with the following parameters: ‘--minimum-mapping-quality 30 --max-reads-per-alignment-start 1000 --min-base-quality-score 20’. The CombineGVCFs module was then used to merge all of the GVCF files, generating a raw population genotype file. The GenotypeGVCFs and SelectVariants modules were used to detect variations and obtain a raw VCF set consisting of only SNPs. The VariantFiltration module was used to remove low-quality SNPs with the following parameters: ‘QD < 2.0 || MQ < 40.0 || FS > 60.0 || SOR >3.0 || MQRankSum < −12.5 || ReadPosRankSum < −8.0’. Finally, VCFtools version 0.1.12b (Danecek et al., 2011) was used to select a subset of high-quality SNPs for subsequent analyses using the following options: ‘--max-missing 0.8 --minQ 30 --min-alleles 2 --max-alleles 2 --maf 0.05’.

For the rice RIL population, SNPs were detected and maps were constructed using methods similar to those described in previous studies (Zhao et al., 2020, 2023). In brief, because the parent ‘Nipponbare’ served as the reference genome, WGS was performed only for the other parent (N22) at ∼36× coverage (producing 13.2 Gb of data). N22 was then aligned to the Nipponbare_V7.0 genome (Kawahara et al., 2013). After filtering out heterozygous and low-quality SNPs, only homozygous SNPs with a depth of at least 5× in N22 were retained. SNP calling was then performed for each RIL. Specifically, barcode-trimmed reads were aligned to the Nipponbare_V7.0 reference genome using BWA–MEM version 0.7.17 (Li and Durbin, 2009). The aligned reads were processed with SAMtools version 1.9 (with the options ‘mpileup -g -Q 20 -q 30’) and BCFtools version 1.6 (with the options ‘call -m -V indels’) (Danecek et al., 2021) to generate a raw VCF file. Candidate SNPs were retained if they met the following criteria: (1) Phred-scaled quality scores (QUAL) ≥ 40; (2) MQ ≥ 20; and (3) homozygosity and consistency with either parental line. Considering the likely spurious SNP calls resulting from sequencing errors or low sequencing depth, a sliding-window approach implemented in the SEG-Map package (Zhao et al., 2010) was used to identify the recombination breakpoints along the chromosomes of each individual. Finally, a dense genetic linkage map was constructed using QTL IciMapping version 4.1 (Meng et al., 2015), with recombination bins serving as genetic markers. Collinearity between the linkage map and the Nipponbare_V7.0 reference genome was assessed by plotting genetic marker coordinates (in centimorgans) against the corresponding physical midpoint positions (in megabases) using the online tool shinyCircos-V2.0 (Wang et al., 2023a).

Phylogenetic tree construction and PCA

A VCF file containing a subset of high-quality SNPs from 127 maize inbred lines was used as the input file of VCF2Dis version 1.47 (https://github.com/BGI-shenzhen/VCF2Dis) to calculate a pairwise distance matrix as described previously (Wang et al., 2023b). The distance matrix was then imported into FastMe version 2.1.6.4 (Lefort et al., 2015) to generate an unrooted NJ phylogenetic tree file. The resulting tree was visualized with iTOL version 6.6 (https://itol.embl.de/). In addition, the VCF file was converted to PLINK format using VCFtools version 0.1.12b (Danecek et al., 2011), and PCA was performed with PLINK version 1.9 (Purcell et al., 2007). The first two eigenvectors were plotted in two dimensions.

Accession numbers

The raw sequencing data reported here have been deposited at the Genome Sequence Archive of the National Genomics Data Center under accession numbers CRA010884, CRA010885, and CRA011752. These data are publicly accessible at https://ngdc.cncb.ac.cn/gsa.

Data and code availability

Custom Perl scripts and the pipeline for de-multiplexing of TAIL-peq sequencing data, read composition analyses, tag identification, and variant calling are available at GitHub (https://github.com/dashengzhao/TAIL-peq/).

Funding

This work was supported by the 10.13039/501100012245 Science and Technology Planning Project of Guangdong Province (2022B0202060002 ), the 10.13039/501100001809 National Natural Science Foundation of China (32300340 , 32172086 ), the R&D program of Shenzhen (KCXFZ20211020164207012 ), and the R&D Program in Key Areas of Guangdong Province (2021B0707010006 ).

Supplemental information

Document S1. Supplemental Figures 1‒15

Document S2. Supplemental Tables 1‒10

Document S3. Article plus supplemental information

Acknowledgments

We thank Lili Dong (China National Center for Bioinformation/Beijing Institute of Genomics, Chinese Academy of Sciences) for assistance with uploading the raw sequencing data. Y.C., S.Z., P.C., et al. are listed as co-inventors on a patent application (CN202211418807.8) filed by the Agricultural Genomics Institute at Shenzhen, Chinese Academy of Agricultural Sciences, describing some of the methods presented here. Y.C. and Y.L. are founders of AgriXY (Shenzhen) Biotechnology Co., Ltd., China.

Author contributions

Conceptualization: Y.C., Q.Q., and S.Z. Methodology: S.Z., Y.W., and Z.Z. Software: S.Z. and Y.W. Validation: Y.W. and P.C. Formal analysis: S.Z. Investigation and resources: P.C., W.L., C.W., H.L., Y.X., and Y.L. Writing – original draft: S.Z. and Y.C. Writing – review & editing: S.Z. and Y.C. Visualization: S.Z. and Y.C. Supervision: Y.C. and Q.Q. Funding acquisition: S.Z. and Y.C. All authors have read and approved the final manuscript.

Published by the Plant Communications Shanghai Editorial Office in association with Cell Press, an imprint of Elsevier Inc., on behalf of CSPB and CEMPS, CAS.

Supplemental information is available at Plant Communications Online.
==== Refs
References

Andrews K.R. Good J.M. Miller M.R. Luikart G. Hohenlohe P.A. Harnessing the power of RADseq for ecological and evolutionary genomics Nat. Rev. Genet. 17 2016 81 92 26729255
Baird N.A. Etter P.D. Atwood T.S. Currey M.C. Shiver A.L. Lewis Z.A. Selker E.U. Cresko W.A. Johnson E.A. Rapid SNP discovery and genetic mapping using sequenced RAD markers PLoS One 3 2008 e3376 18852878
Bernardo A. St Amand P. Le H.Q. Su Z. Bai G. Multiplex restriction amplicon sequencing: a novel next-generation sequencing-based marker platform for high-throughput genotyping Plant Biotechnol. J. 18 2020 254 265 31199572
Bhat J.A. Ali S. Salgotra R.K. Mir Z.A. Dutta S. Jadon V. Tyagi A. Mushtaq M. Jain N. Singh P.K. Genomic selection in the era of next generation sequencing for complex traits in plant breeding Front. Genet. 7 2016 221 28083016
Bradbury L.M.T. Fitzgerald T.L. Henry R.J. Jin Q. Waters D.L.E. The gene for fragrance in rice Plant Biotechnol. J. 3 2005 363 370 17129318
Cai X.L. Wang Z.Y. Xing Y.Y. Zhang J.L. Hong M.M. Aberrant splicing of intron 1 leads to the heterogeneous 5' UTR and decreased expression of waxy gene in rice cultivars of intermediate amylose content Plant J. 14 1998 459 465 9670561
Chen S. Zhou Y. Chen Y. Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor Bioinformatics 34 2018 884 890 29126246
Danecek P. Auton A. Abecasis G. Albers C.A. Banks E. DePristo M.A. Handsaker R.E. Lunter G. Marth G.T. Sherry S.T. The variant call format and VCFtools Bioinformatics 27 2011 2156 2158 21653522
Danecek P. Bonfield J.K. Liddle J. Marshall J. Ohan V. Pollard M.O. Whitwham A. Keane T. McCarthy S.A. Davies R.M. Twelve years of SAMtools and BCFtools GigaScience 10 2021 giab008 33590861
Davey J.W. Hohenlohe P.A. Etter P.D. Boone J.Q. Catchen J.M. Blaxter M.L. Genome-wide genetic marker discovery and genotyping using next-generation sequencing Nat. Rev. Genet. 12 2011 499 510 21681211
Du H. Yu Y. Ma Y. Gao Q. Cao Y. Chen Z. Ma B. Qi M. Li Y. Zhao X. Sequencing and de novo assembly of a near complete indica rice genome Nat. Commun. 8 2017 15324 28469237
Elshire R.J. Glaubitz J.C. Sun Q. Poland J.A. Kawamoto K. Buckler E.S. Mitchell S.E. A Robust, Simple Genotyping-by-Sequencing (GBS) Approach for High Diversity Species PLoS One 6 2011 e19379 21573248
Gao Z. Zeng D. Cui X. Zhou Y. Yan M. Huang D. Li J. Qian Q. Map-based cloning of the ALK gene, which controls the gelatinization temperature of rice Sci. China, Ser. A C. 46 2003 661 668
Guo Z. Yang Q. Huang F. Zheng H. Sang Z. Xu Y. Zhang C. Wu K. Tao J. Prasanna B.M. Development of high-resolution multiple-SNP arrays for genetic analyses and molecular breeding through genotyping by target sequencing and liquid chip Plant Commun. 2 2021 100230 34778746
Hodges E. Rooks M. Xuan Z. Bhattacharjee A. Benjamin Gordon D. Brizuela L. Richard McCombie W. Hannon G.J. Hybrid selection of discrete genomic intervals on custom-designed microarrays for massively parallel sequencing Nat. Protoc. 4 2009 960 974 19478811
Hu Y. Colantonio V. Müller B.S.F. Leach K.A. Nanni A. Finegan C. Wang B. Baseggio M. Newton C.J. Juhl E.M. Genome assembly and population genomic analysis provide insights into the evolution of modern sweet corn Nat. Commun. 12 2021 1227 33623026
Huang X. Feng Q. Qian Q. Zhao Q. Wang L. Wang A. Guan J. Fan D. Weng Q. Huang T. High-throughput genotyping by whole-genome resequencing Genome Res. 19 2009 1068 1076 19420380
Huang X. Zhao Y. Wei X. Li C. Wang A. Zhao Q. Li W. Guo Y. Deng L. Zhu C. Genome-wide association study of flowering time and grain yield traits in a worldwide collection of rice germplasm Nat. Genet. 44 2012 32 39
Jiao Y. Peluso P. Shi J. Liang T. Stitzer M.C. Wang B. Campbell M.S. Stein J.C. Wei X. Chin C.S. Improved maize reference genome with single-molecule technologies Nature 546 2017 524 527 28605751
Kawahara Y. de la Bastide M. Hamilton J.P. Kanamori H. McCombie W.R. Ouyang S. Schwartz D.C. Tanaka T. Wu J. Zhou S. Improvement of the Oryza sativa Nipponbare reference genome using next generation sequence and optical map data Rice 6 2013 4 24280374
Kozarewa I. Armisen J. Gardner A.F. Slatko B.E. Hendrickson C.L. Overview of Target Enrichment Strategies Curr. Protoc. Mol. Biol. 112 2015 7.21.1 7.21.23 21-27.21.23
Kwon S.J. Hong S.W. Son J.H. Lee J.K. Cha Y.S. Eun M.Y. Kim N.S. CACTA and MITE transposon distributions on a genetic map of rice using F15 RILs derived from Milyang 23 and Gihobyeo hybrids Mol. Cells 21 2006 360 366 16819298
Lefort V. Desper R. Gascuel O. FastME 2.0: A Comprehensive, Accurate, and Fast Distance-Based Phylogeny Inference Program Mol. Biol. Evol. 32 2015 2798 2800 26130081
Li H. Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform Bioinformatics 25 2009 1754 1760 19451168
Li H. Handsaker B. Wysoker A. Fennell T. Ruan J. Homer N. Marth G. Abecasis G. Durbin R. 1000 Genome Project Data Processing Subgroup The Sequence Alignment/Map format and SAMtools Bioinformatics 25 2009 2078 2079 19505943
Liu K. Goodman M. Muse S. Smith J.S. Buckler E. Doebley J. Genetic structure and diversity among maize inbred lines as inferred from DNA microsatellites Genetics 165 2003 2117 2128 14704191
Liu Y.G. Chen Y. High-efficiency thermal asymmetric interlaced PCR for amplification of unknown flanking sequences Biotechniques 43 2007 649 656 18072594
Liu Y.G. Whittier R.F. Thermal asymmetric interlaced PCR: automatable amplification and sequencing of insert end fragments from P1 and YAC clones for chromosome walking Genomics 25 1995 674 681 7759102
Lu Q. Huang L. Liu H. Garg V. Gangurde S.S. Li H. Chitikineni A. Guo D. Pandey M.K. Li S. A genomic variation map provides insights into peanut diversity in China and associations with 28 agronomic traits Nat. Genet. 56 2024 530 540 10.1038/s41588-024-01660-7 In press 38378864
Lu Y. Chen X. Yu H. Zhang C. Xue Y. Zhang Q. Wang H. Haplotype-resolved genome assembly of Phanera championii reveals molecular mechanisms of flavonoid synthesis and adaptive evolution Plant J. 118 2024 488 505 10.1111/tpj.16620 38173092
Markoulatos P. Siafakas N. Moncany M. Multiplex polymerase chain reaction: a practical approach J. Clin. Lab. Anal. 16 2002 47 51 11835531
Marks R.A. Hotaling S. Frandsen P.B. VanBuren R. Representation and participation across 20 years of plant genome sequencing Nat. Plants 7 2021 1571 1578 34845350
McKenna A. Hanna M. Banks E. Sivachenko A. Cibulskis K. Kernytsky A. Garimella K. Altshuler D. Gabriel S. Daly M. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data Genome Res. 20 2010 1297 1303 20644199
Meng L. Li H. Zhang L. Wang J. QTL IciMapping: Integrated software for genetic linkage map construction and quantitative trait locus mapping in biparental populations Crop J. 3 2015 269 283
Munyengwa N. Le Guen V. Bille H.N. Souza L.M. Clément-Demange A. Mournet P. Masson A. Soumahoro M. Kouassi D. Cros D. Optimizing imputation of marker data from genotyping-by-sequencing (GBS) for genomic selection in non-model species: Rubber tree (Hevea brasiliensis) as a case study Genomics 113 2021 655 668 33508443
Nemoto Y. Nonoue Y. Yano M. Izawa T. Hd1,a CONSTANS ortholog in rice, functions as an Ehd1 repressor through interaction with monocot-specific CCT-domain protein Ghd7 Plant J. 86 2016 221 233 26991872
Nguyen N.N. Kim M. Jung J.K. Shim E.J. Chung S.M. Park Y. Lee G.P. Sim S.C. Genome-wide SNP discovery and core marker sets for assessment of genetic variations in cultivated pumpkin (Cucurbita spp.) Hortic. Res. 7 2020 121 32821404
Purcell S. Neale B. Todd-Brown K. Thomas L. Ferreira M.A.R. Bender D. Maller J. Sklar P. de Bakker P.I.W. Daly M.J. PLINK: A tool set for whole-genome association and population-based linkage analyses Am. J. Hum. Genet. 81 2007 559 575 17701901
Rasheed A. Hao Y. Xia X. Khan A. Xu Y. Varshney R.K. He Z. Crop Breeding Chips and Genotyping Platforms: Progress, Challenges, and Perspectives Mol. Plant 10 2017 1047 1064 28669791
Saghai-Maroof M.A. Soliman K.M. Jorgensen R.A. Allard R.W. Ribosomal DNA spacer-length polymorphisms in barley: mendelian inheritance, chromosomal location, and population dynamics Proc. Natl. Acad. Sci. USA 81 1984 8014 8018 6096873
Shang L. He W. Wang T. Yang Y. Xu Q. Zhao X. Yang L. Zhang H. Li X. Lv Y. A complete assembly of the rice Nipponbare reference genome Mol. Plant 16 2023 1232 1236 37553831
Shen Y. Zhou G. Liang C. Tian Z. Omics-based interdisciplinarity is accelerating plant breeding Curr. Opin. Plant Biol. 66 2022 102167 35016139
Sims D. Sudbery I. Ilott N.E. Heger A. Ponting C.P. Sequencing depth and coverage: key considerations in genomic analyses Nat. Rev. Genet. 15 2014 121 132 24434847
Sun M. Yao C. Shu Q. He Y. Chen G. Yang G. Xu S. Liu Y. Xue Z. Wu J. Telomere-to-telomere pear (Pyrus pyrifolia) reference genome reveals segmental and whole genome duplication driving genome evolution Hortic. Res. 10 2023 uhad201 38023478
Sun Y. Shang L. Zhu Q.H. Fan L. Guo L. Twenty years of plant genome sequencing: achievements and challenges Trends Plant Sci. 27 2022 391 401 34782248
Wallace J.G. Mitchell S.E. Genotyping-by-Sequencing Curr. Protoc. Plant Biol. 2 2017 64 77 31725977
Wang Y. Jia L. Ge T. Dong Y. Zhang X. Zhou Z. Luo X. Li Y. Yao W. shinyCircos-V2.0: Leveraging the creation of Circos plot with enhanced usability and advanced features iMeta 2 2023 e109 38868422
Wang Y. Zhao S. Chen P. Liu Y. Ma Z. Malik W.A. Zhu Z. Peng Z. Lu H. Chen Y. Genetic Diversity and Population Structure Analysis of Hollyhock (Alcea rosea Cavan) Using High-Throughput Sequencing Horticulturae 9 2023 662
Watowich M.M. Chiou K.L. Graves B. Montague M.J. Brent L.J.N. Higham J.P. Horvath J.E. Lu A. Martinez M.I. Platt M.L. Best practices for genotype imputation from low-coverage sequencing data in natural populations Mol. Ecol. Resour. 2023 10.1111/1755-0998.13854 In press
Wu W. Zheng X.M. Lu G. Zhong Z. Gao H. Chen L. Wu C. Wang H.J. Wang Q. Zhou K. Association of functional nucleotide polymorphisms at DTH2 with the northward expansion of rice cultivation in Asia Proc. Natl. Acad. Sci. USA 110 2013 2775 2780 23388640
Wurschum T. Mapping QTL for agronomic traits in breeding populations Theor. Appl. Genet. 125 2012 201 210 22614179
Xie W. Feng Q. Yu H. Huang X. Zhao Q. Xing Y. Yu S. Han B. Zhang Q. Parent-independent genotyping for constructing an ultrahigh-density linkage map based on population sequencing Proc. Natl. Acad. Sci. USA 107 2010 10578 10583 20498060
Xu Y. Zhang X. Li H. Zheng H. Zhang J. Olsen M.S. Varshney R.K. Prasanna B.M. Qian Q. Smart breeding driven by big data, artificial intelligence, and integrated genomic-enviromic prediction Mol. Plant 15 2022 1664 1695 36081348
Yang Y. Wu Z. Wu Z. Li T. Shen Z. Zhou X. Wu X. Li G. Zhang Y. A near-complete assembly of asparagus bean provides insights into anthocyanin accumulation in pods Plant Biotechnol. J. 21 2023 2473 2489 37558431
Yao F.Y. Xu C.G. Yu S.B. Li J.X. Gao Y.J. Li X.H. Zhang Q. Mapping and genetic analysis of two fertility restorer loci in the wild-abortive cytoplasmic male sterility system of rice (Oryza sativa L.) Euphytica 98 1997 183 187
Zhang G. Lu Y. Bharaj T.S. Virmani S.S. Huang N. Mapping of the Rf-3 nuclear fertility-restoring gene for WA cytoplasmic male sterility in rice using RAPD and RFLP markers Theor. Appl. Genet. 94 1997 27 33 19352741
Zhang H. He Q. Xing L. Wang R. Wang Y. Liu Y. Zhou Q. Li X. Jia Z. Liu Z. The haplotype-resolved genome assembly of autotetraploid rhubarb Rheum officinale provides insights into its genome evolution and massive accumulation of anthraquinones Plant Commun. 5 2024 100677 37634079
Zhang H. Wang Y. Deng C. Zhao S. Zhang P. Feng J. Huang W. Kang S. Qian Q. Xiong G. High-quality genome assembly of Huazhan and Tianfeng, the parents of an elite rice hybrid Tian-you-hua-zhan Sci. China Life Sci. 65 2022 398 411 34251582
Zhao Q. Huang X. Lin Z. Han B. SEG-Map: A Novel Software for Genotype Calling and Genetic Map Construction from Next-generation Sequencing Rice 3 2010 98 102
Zhao S. Li X. Song J. Li H. Zhao X. Zhang P. Li Z. Tian Z. Lv M. Deng C. Genetic dissection of maize plant architecture using a novel nested association mapping population Plant Genome 15 2022 e20179 34859966
Zhao S. Zhang C. Mu J. Zhang H. Yao W. Ding X. Ding J. Chang Y. All-in-one sequencing: an improved library preparation method for cost-effective and high-throughput next-generation sequencing Plant Methods 16 2020 74 32489396
Zhao S. Zhang C. Wang L. Luo M. Zhang P. Wang Y. Malik W.A. Wang Y. Chen P. Qiu X. A prolific and robust whole-genome genotyping method using PCR amplification via primer-template mismatched annealing J. Integr. Plant Biol. 65 2023 633 645 36269601
Zhou B. Qu S. Liu G. Dolan M. Sakai H. Lu G. Bellizzi M. Wang G.L. The eight amino-acid differences within three leucine-rich repeats between Pi2 and Piz-t resistance proteins determine the resistance specificity to Magnaporthe grisea Mol. Plant Microbe. In. 19 2006 1216 1228
Zou M. Xia Z. Hyper-seq: A novel, effective, and flexible marker-assisted selection and genotyping approach Innovation 3 2022 100254 35602119
