
==== Front
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
GigaScience Press Sha Tin, New Territories, Hong Kong SAR

DRR-202405-01
134
10.46471/gigabyte.134
https://doi.org/10.1101/2024.08.16.608262
Data Release
Genetics and Genomics
Animal Genetics
Animal Physiology
High-speed whole-genome sequencing of a Whippet: Rapid chromosome-level assembly and annotation of an extremely fast dog’s genome
M. Nebenführ et al.
High-speed whole-genome sequencing of a Whippet: Rapid chromosome-level assembly and annotation of an extremely fast dog’s genome
https://orcid.org/0000-0001-8802-2105
Nebenführ Marcel 1 2 3 Formal analysis Writing - original draft Visualization Writing - review editing * †
https://orcid.org/0009-0000-6275-7752
Prochotta David 1 2 3 Formal analysis Writing - original draft Visualization Writing - review editing †
https://orcid.org/0009-0008-9826-902X
Ben Hamadou Alexander 1 3 Methodology Writing - review editing
https://orcid.org/0000-0002-9394-1904
Janke Axel 1 2 3 Funding acquisition Supervision Writing - review editing
https://orcid.org/0009-0005-0763-4780
Gerheim Charlotte 1 3 Methodology Writing - review editing
Betz Christian 4 Methodology Writing - review editing
https://orcid.org/0000-0003-4993-1378
Greve Carola 1 3 Conceptualization Writing - original draft Writing - review editing
https://orcid.org/0000-0001-9317-0022
Bolz Hanno Jörn 4 Conceptualization Writing - original draft Writing - review editing *
1 Senckenberg Biodiversity and Climate Research Centre (BiK-F), Frankfurt am Main, Germany
2 Institute for Ecology, Evolution, and Diversity, Goethe University, Frankfurt am Main, Germany
3 LOEWE-Centre for Translational Biodiversity Genomics (TBG), Frankfurt am Main, Germany
4 Bioscientia Human Genetics, Institute for Medical Diagnostics GmbH, Ingelheim, Germany
* Corresponding authors. E-mail: marcel.nebenfuehr@senckenberg.de; hanno.bolz@bioscientia.de
† Contributed equally.

13 9 2024
2024
2024 gigabyte13422 5 2024
09 9 2024
© The Author(s) 2024.
2024
https://creativecommons.org/licenses/by/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
The time required for genome sequencing and de novo assembly depends on the interaction between laboratory work, sequencing capacity, and the bioinformatics workflow, often constrained by external sequencing services. Bringing together academic biodiversity institutes and a medical diagnostics company with extensive sequencing capabilities, we aimed at generating a high-quality mammalian de novo genome in minimal time. We present the first chromosome-level genome assembly of the Whippet, using PacBio long-read high-fidelity sequencing and reference-guided scaffolding. The final assembly has a contig N50 of 55 Mbp and a scaffold N50 of 65.7 Mbp. The total assembly length is 2.47 Gbp, of which 2.43 Gpb were scaffolded into 39 chromosome-length scaffolds. Annotation using mammalian genomes and transcriptome data yielded 28,383 transcripts, 90.9% complete BUSCO genes, and identified 36.5% repeat content. Sequencing, assembling, and scaffolding the chromosome-level genome of the Whippet took less than a week, adding another high-quality reference genome to the available sequences of domestic dog breeds.

Hessen State Ministry of Higher Education, Research and the Arts LOEWE/1/10/519/03/03.001(0014)/52 The contribution of the Centre for Translational Biodiversity Genomics (LOEWE-TBG) to the study was funded by the Hessen State Ministry of Higher Education, Research and the Arts (LOEWE/1/10/519/03/03.001(0014)/52).
==== Body
pmcData description

Background information

Although recent advances in sequencing technologies made genomics more accessible to a broader scientific community, sequencing and assembling large eukaryotic genomes of non-model organisms remains challenging. Even with access to samples of sufficient quality, limited funding for sequencing, insufficient computing power, lack of sequencing technology and capacity, and the need for outsourcing can significantly increase turnaround times.

In response to the accelerating global biodiversity loss, national and international initiatives expanded genomic references for biodiversity research and conservation. The use of reference genomes in population genomics facilitates the characterization of genetic diversity and adaptation through local variant enrichment, providing a basis for biodiversity assessment, conservation, and restoration [1–3].

Furthermore, non-human genomes have become increasingly helpful for the interpretation of human genetic variants and their relevance for disease [4, 5]; this particularly applies to de novo-assembled long-read human genomes that, in contrast to short-read genomes, allow for the identification of complex (structural) variants that would otherwise escape detection. In both biodiversity and conservation research, as well as in medical genetics, the timely generation and assembly of genomic data is of utmost importance.

In this work, an academic biodiversity institute joined forces with a medical diagnostic company with extensive sequencing capacity and experience to generate a rapid chromosome-level de novo genome of the Whippet. We demonstrated that streamlined laboratory and bioinformatics workflows with PacBio high-fidelity (HiFi) long-read whole-genome sequencing (LR-WGS) enable the generation of a high-quality reference genome within one week (Figure 1). Besides this proof-of-concept study of the rapid Whippet genome, our collaboration includes continuous LR-WGS of de novo genomes of various endangered species (including non-vertebrates and plants). Our study may serve as a paradigm for such cooperations applicable to a wide range of human and non-human genome projects, from biodiversity research to domestic animal and agricultural research.

Figure 1. Project timeline with day-by-day progress description and time requirements.

The blue line represents the contribution of the biodiversity research centre (TBG) at each step, and the red line represents the contributions made by the medical diagnostics company (Bioscientia).

Sampling, DNA extraction, and sequencing

High molecular weight genomic DNA was extracted from the peripheral blood leukocytes of a three-year-old male dog (Canis lupus familiaris, NCBI:txid9615), a Whippet, using the PacBio Nanobind CBB kit (Pacific Biosciences, Menlo Park, CA). Blood was taken during a routine veterinary procedure, collected in EDTA-coated vials, and frozen at −20 °C. DNA concentration and DNA fragment lengths were evaluated using the Qubit dsDNA BR Assay kit on the Qubit Fluorometer (Thermo Fisher Scientific, Waltham, MA) and the Genomic DNA Screen Tape on the Agilent 4150 TapeStation system (Agilent Technologies, Santa Clara, CA), respectively. Two SMRTbell libraries were prepared according to the instructions in the SMRTbell Express Prep Kit v3.0. The final concentrations were 68 ng/μl and 76 ng/μl, with a total input of approximately 10 μg of sheared DNA per library. Annealing of sequencing primers, binding of sequencing polymerase, and purification of polymerase-bound SMRTbell complexes were performed using the Revio polymerase kit (PacBio, Menlo Park, CA, USA). The loading concentration for sequencing was 250 pM.

Genome assembly and polishing

The two sequencing runs on a PacBio Revio® instrument of the Whippet yielded a total of ∼128 Gbp of sequence data, with an average subread N50 of ∼17.8 kbp. We assembled the genome using Hifiasm v0.18.8-r525 (RRID:SCR_021069) [6] with default settings and scanned the initial genome assembly for contamination using FCS-GX v0.4.0 [7].

We polished the raw assembly three times using Inspector v1.0.1 [8], which includes Flye v2.9.3 (RRID:SCR_017016) [9] for structural error correction.

After polishing, we performed reference-guided scaffolding using RagTag v2.1.0 [10] with default settings. For that, the NCBI reference genome for dogs was chosen as a reference, which is the German Shepherd dog genome (GCA_011100685.1). Next, we used TGS-GapCloser v1.2.1 (RRID:SCR_017633) [11] to close assembly gaps using three rounds of gap-filling and the Racon v1.0.5 (RRID:SCR_017642) [12] module for polishing the assembly gaps.

Assembly quality control

We evaluated both the raw assembly and the final assembly with Merqury v1.3 (RRID:SCR_022964) [13] in combination with Meryl v1.4.1, as well as Inspector v1.0.1.

We calculated assembly contiguity statistics of the final genome using Inspector. We then performed a gene set completeness analysis of the Whippet genome and other available dog genomes for comparison using BUSCO v5.4.22 (RRID:SCR_015008) [14] with the provided Carnivora orthologous genes database (carnivora_odb10).

We mapped the reads back to the genome using minimap2 v2.28 (RRID:SCR_018550) [15] with the ‘-ax map-hifi’ option to output a SAM file. The SAM file was sorted and converted to a BAM file with samtools v1.20 (RRID:SCR_002105) [16]. We then marked duplicate sequences with sambamba v1.0.1 (RRID:SCR_024328) [17] and analyzed mapping statistics with QualiMap v2.2.1 (RRID:SCR_001209) [18].

The initial assembly, based on the HiFi long reads only, already recovered 36 (of 39) chromosome-sized scaffolds. This further demonstrates how effective and therefore invaluable accurate long-read sequencing has become in mammalian genomics.

The final polished, scaffolded, and gap-closed assembly consists of 148 scaffolds with a total length of 2.47 Gbp, while 2.43 Gbp were placed into 39 chromosome-sized scaffolds (Figure 2, A and B; Table 1). Gene completeness analysis of the genome based on BUSCO’s Carnivora dataset identified 14,170 complete single-copy orthologous sequences, corresponding to 97.72% completeness and 77 (0.53%) missing genes (Figure 2A).

Figure 2. (A) Snailplot based on the polished and reference-scaffolded assembly showing BUSCO gene completeness results and basic assembly statistics. (B) Whole-genome synteny between the chromosome-level assembly of the German Shepherd and our chromosome-level assembly of the Whippet.

Table 1 Basic assembly statistics, Inspector and Merqury quality values of the raw and final Whippet assembly.

	Raw assembly	Final assembly	
Total length (bp)	2,472,523,283	2,472,512,883	
No. of contigs/scaffolds	172	148	
N50	54,692,731	65,741,809	
L50	17	15	
Longest contig	123,386,781	124,133,413	
Inspector QV	50.2	53.5	
Inspector mapping rate (%)	100	100	
Merqury QV	65.6	68.1	
Merqury completeness	97.87%	97.87%	

Both Inspector and Merqury, coupled with the high BUSCO score, indicated a highly contiguous, accurate, and complete genome. In addition, available dog assemblies were downloaded to compare the BUSCO completeness across published reference genomes (Figure 3, Table 2). In addition, the qualimap report for the resulting BAM file showed 99.98% of mapped reads (8,332,294), a mean mapping quality of 49.7, a mean coverage of 55.1X, and an error rate of 0.0057.

Figure 3. Comparison of BUSCO completeness statistics based on the carnivora database between our final Whippet assembly and annotation, and other available dog assemblies.

Complete single-copy genes are shaded light blue, and complete duplicated sequences are shaded dark blue; fragmented genes are shaded yellow, and missing sequences are shaded red. The numbers of complete single copy (S), complete duplicated (D), fragmented (F), and missing genes (M) for the respective genome are shown in each column. The total number of genes in the BUSCO carnivora library is denoted as n.

Table 2 Available genome data used for comparison.

Species	Accession number	
German Shepherd (Canis lupus familiaris)	GCA_011100685.1	
Labrador Retriever (Canis lupus familiaris)	GCA_012044875.1	
Basenji (Canis lupus familiaris)	GCA_013276365.2	
Cairn Terrier (Canis lupus familiaris)	GCA_031010295.1	
Bernese Mountain Dog (Canis lupus familiaris)	GCA_031010765.1	
Boxer (Canis lupus familiaris)	GCF_000002285.5	
Cat (Felis catus)	GCF_018350175.1	
Dingo (Canis lupus dingo)	GCF_003254725.2	
Mouse (Mus musculus)	GCF_000001635.27	
Human (Homo sapiens)	GCF_009914755.1	

Heterozygosity

To calculate genome-wide heterozygosity in the Whippet genome, we first mapped the reads used for assembly to the genome and marked duplicates using sambamba v1.0.0, then counted base-depth at all sites using sambamba. We then estimated the Site Frequency Spectrum (SFS) with ANGSD v0.940 (RRID:SCR_021865) [19] and used the output files to run realSFS (RRID:SCR_002493) with 200 bootstrap replicates to calculate the folded SFS. From this, we calculated the heterozygosity by dividing the heterozygous sites by the sum of the homozygous and heterozygous sites. In line with other dog genomes, the heterozygosity was 0.09%. Our heterozygosity analysis delivers standard baseline data for a de novo sequenced genome and allows a first glimpse into the genetic diversity of Whippets. However, it cannot be representative of Whippets in general.

Genome annotation

Repeat annotation

To annotate repetitive regions in the Whippet genome, we used RepeatModeler v2.1 (RRID:SCR_015027) [20] to create a de novo repeat library for our assembly. Next, we used RepeatMasker v4.1.6 (RRID:SCR_012954) [21] to hard-mask repeats based on the modeled repeats. Our analyses identified 36.5% of repeats in the genome, of which the majority consisted of long interspersed nuclear elements (LINEs) (20.32%) and long terminal repeat (LTR) elements (3.68%). In addition, 5.37% of unclassified elements were identified (Table 3).

Table 3 Repeat content of the Whippet genome assembly. Class: class of the repetitive regions. Count: number of occurrences of the repetitive region. bpMasked: number of base pairs masked; %masked: percentage of base pairs masked. LINE: Long Interspersed Nuclear Elements (including retroposons); LTR: Long Terminal Repeat elements (including retroposons); SINE: Short Interspersed Nuclear Elements; RC: Rolling Circle. In total, 902,477,158 bp were masked, corresponding to 36.5% of the genome.

Class	Count	bpMasked	%masked	
SINEs	478,438	65,159,195	2.64	
LINEs	1,796,370	502,326,946	20.32	
LTR	354,096	90,945,042	3.68	
DNA transposons	300,163	49,080,228	1.99	
Rolling-circles	1,568	78,131	0.00	
Unclassified	525,739	132,765,950	5.37	
Small RNA	61,933	5,023,647	0.2	
Satellites	14,007	2,268,618	0.09	
Simple repeats	870,502	48,791,971	1.97	
Low complexity	109,450	6,037,430	0.24	

Gene annotation

Gene annotation was performed on the unmasked assembly using GeMoMa v1.9 (RRID:SCR_017646) [22]. Therefore, available annotated high-quality genomes of dogs and other mammals were used as references to identify genes (Table 2). The final gene annotation resulted in 28,383 transcripts. In addition, our BUSCO analysis identified 90.9% complete BUSCOs, suggesting a high annotation completeness (Figure 3).

Myostatin

Since a founder mutation in myostatin (MSTN) has been reported in Whippet dogs [23], we analyzed our Whippet genome for this mutation by comparing its MSTN sequences to the wild-type references of Canis lupus familiaris (AY367768) and found no mutation in the sequence (Figure 4). We first identified the position of the gene within the genome with MMseqs2 (RRID:SCR_022962) [24] and then visualized the region in IGV (RRID:SCR_011793) [25] to check for mutations in the aligned reads. Hence, the analysis of our Whippet genome’s MSTN sequence primarily exemplifies a possible application of available high-quality genomes, especially for breeding purposes, but is not representative of Whippets in general.

Figure 4. Read alignment of the myostatin gene, MSTN, in the sequenced Whippet genome.

The black box indicates the MSTN sequence of interest, as used by Mosher et al. [23].

Conclusion

In our proof-of-concept study, we show that teaming up a medical diagnostics company with a biodiversity research institute may deliver extremely rapid de novo-assembled HiFi long-read genomes. This was possible through close and streamlined time management and collaboration, including all required participants for a genome project, namely a veterinarian, laboratory facilities, the sequencing facility, and the bioinformatics unit.

Acknowledgements

We thank Kristina Grund, Kleintierpraxis Auringen, Germany, for veterinary care, and Ursula Wollscheid, Inge Lischewski, Lea Arndt and Anna Linck for technical assistance. We also thank Christian Decker and Sebastian Görges for bioinformatics support.

Data availability

All raw data generated in this study are accessible at GenBank under BioProject PRJNA1114051. Annotation, results files, and other data are available in the GigaDB repository [26].

Abbreviations

HiFi, High Fidelity; LINE, long interspersed nuclear element; LR-WGS, long-read whole-genome sequencing; LTR, long terminal repeat; MSTN, myostatin; SFS, Site Frequency Spectrum.

Declarations

Ethics approval and consent to participate

Not applicable.

Competing interests

CB and HJB are employees of Bioscientia, which is part of a publicly traded diagnostic company. The authors declare that they have no competing interests.

Author contributions

ABH and ChG performed the DNA extraction and the library preparation. CB supervised genome sequencing and data transfer. DP and MN assembled and analyzed the genomes, and conducted the downstream analyses. CG, MN, DP, and HJB jointly supervised the project and wrote the manuscript with input from ABH, AJ, ChG, and CB. All authors read and approved the final manuscript before submission.

Funding

The contribution of the Centre for Translational Biodiversity Genomics (LOEWE-TBG) to the study was funded by the Hessen State Ministry of Higher Education, Research and the Arts (LOEWE/1/10/519/03/03.001(0014)/52).

GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Article Submission
Nebenführ Marcel Mr Author
21 5 2024
22 5 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Assign Handling Editor
Edmunds Scott Dr Editor in Chief
22 5 2024
22 5 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Editor Assess MS
Zhang Hongfang Dr Handling Editor
22 5 2024
23 5 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Curator Assess MS
Fan Yannan Ms Curator
03 6 2024
13 6 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Review MS
Lan Tianming Mr Reviewer 1
10 7 2024
10 7 2024
Reviewer name and names of any other individual's who aided in reviewer	Tianming Lan	
Do you understand and agree to our policy of having open and named reviews, and having your review included with the published papers. (If no, please inform the editor that you cannot review this manuscript.)	Yes	
Is the language of sufficient quality?	Yes	
Please add additional comments on language quality to clarify if needed		
Are all data available and do they match the descriptions in the paper?	Yes	
Additional Comments		
Are the data and metadata consistent with relevant minimum information or reporting standards? See GigaDB checklists for examples <a href="http://gigadb.org/site/guide" target="_blank">http://gigadb.org/site/guide</a>	Yes	
Additional Comments		
Is the data acquisition clear, complete and methodologically sound?	Yes	
Additional Comments		
Is there sufficient detail in the methods and data-processing steps to allow reproduction?	Yes	
Additional Comments		
Is there sufficient data validation and statistical analyses of data quality?	Not my area of expertise	
Additional Comments		
Is the validation suitable for this type of data?	Yes	
Additional Comments		
Is there sufficient information for others to reuse this dataset or integrate it with other data?	Yes	
Additional Comments		
Any Additional Overall Comments to the Author	The authors provided an example of High-speed strategy for whole-genome sequencing, genome assembly and annotation for species and take an example with the Whippet dog. This is a very novel idea under the genomic era with plummeting sequencing cost, fast accumulated sequencing data but shortage of computing resources. The authors also provide a very high-quality reference genome for the Whippet dog species with very good contiguity, accuracy and completeness. However, I have several concerns need the authors to further consider before it could be published at the journal of Gigatyte. Q1. There are too many keywords. Can the authors reduce a few? Biodiversity conservation, Comparative genomics, and evolutionary biology does not make sense in this manuscript. Q2. The authors performed reference-guided scaffolding analysis with the German Shepherd dog genome (GCA_011100685.1) as reference. Better if the authors explain why they selected this genome as the reference as there are several published dog genomes? Q3.The part of Heterozygosity make no sense to this manuscript unless there is a reasonable connection with other parts, because the dog is not a threatened species and also not a very special breed facing extensive inbreeding abd accumulation of deleterious mutations? Q4. The part of Myostatin doesn’t make sense to me, as I have read the paper the author cited and found that not all Whippet have this mutation? They sequenced 22 individuals, and 4 individuals are homozygous (-/-), 5 are heterozygous (mh/+) and the rest are homozygous (+/+). So you can always have a result by checking this mutation, but make no sense. Furthermore, one individual can hardly represent a species or a population? At the beginning of this paragraph, please change “Since” to “Since”. Q5. I think the most important find in this manuscript is how the authors finished a high-quality genome within a very short-term working. I suggest the authors remove the descriptions of Heterozygosity and Myostatin, but added a paragraph to tell readers the basic needs or standards for such a short-term work for genome assembly for a genome of something like dog. Just a suggestion, but I think would be better to improve the manuscript.	
Recommendation	Major Revision	

GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Review MS
Wang Xiaobo Dr Reviewer 2
16 7 2024
21 7 2024
Reviewer name and names of any other individual's who aided in reviewer	Xiaobo Wang	
Do you understand and agree to our policy of having open and named reviews, and having your review included with the published papers. (If no, please inform the editor that you cannot review this manuscript.)	Yes	
Is the language of sufficient quality?	Yes	
Please add additional comments on language quality to clarify if needed		
Are all data available and do they match the descriptions in the paper?	No	
Additional Comments	No link to the relevant data in GigaDB was provided.	
Are the data and metadata consistent with relevant minimum information or reporting standards? See GigaDB checklists for examples <a href="http://gigadb.org/site/guide" target="_blank">http://gigadb.org/site/guide</a>	No	
Additional Comments	No link to the relevant data in GigaDB was provided.	
Is the data acquisition clear, complete and methodologically sound?	Yes	
Additional Comments		
Is there sufficient detail in the methods and data-processing steps to allow reproduction?	Yes	
Additional Comments		
Is there sufficient data validation and statistical analyses of data quality?	Yes	
Additional Comments		
Is the validation suitable for this type of data?	Yes	
Additional Comments		
Is there sufficient information for others to reuse this dataset or integrate it with other data?	Yes	
Additional Comments		
Any Additional Overall Comments to the Author	This study outlines an approach to expedite the sequencing and de novo assembly of genomes by leveraging collaboration between academic biodiversity institutes and a medical diagnostics company with advanced sequencing capabilities. The primary focus was on generating a high-quality de novo genome of the Whippet, a fast dog breed, within an accelerated timeframe. Below are some specific comments I would like to highlight. 1. The authors mentioned the use of QUAST and QualiMap software tools to assess the genome of the Whippet; however, the corresponding results were not presented in the manuscript. 2. The authors' reliance solely on mammalian protein sequences for homology annotation means that unique genes specific to the Whippet remain unannotated. The discrepancy of approximately 7% between the completeness assessments of the gene set and the genome via BUSCO further underscores the incomplete nature of the gene set. To address this, I recommend integrating transcriptome data, at the very least, to incorporate de novo annotation results. This addition should enhance the comprehensiveness and accuracy of gene annotations for the Whippet genome. 3. The authors claim the absence of reported mutations in the Mstn gene but have not provided corroborating evidence, such as read alignment results from the genomic region, to verify that this is not due to assembly errors. 4. If feasible, I propose integrating second-generation sequencing to further polish the genome and elevate its quality.	
Recommendation	Minor Revision	

GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Editor Decision
Zhang Hongfang Dr Handling Editor
21 7 2024
22 7 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Major Revision
Nebenführ Marcel Mr Author
06 8 2024
07 8 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Assess Revision
Zhang Hongfang Dr Handling Editor
07 8 2024
08 8 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Re-Review MS
Lan Tianming Mr Reviewer 1
08 8 2024
08 8 2024
	Indicate in the comments box below whether you are happy with the changes made or if the manuscript is unacceptable.	
Comments on revised manuscript	I saw the authors made some changes which has improved the mauscript, but I still have some minor comments which I think the authors should be further consider before it could be published. 1. I recommend the authors further remove key words from the current version, so redundent with the title, at least delete the biodiversity conservation, because this manuscript has nothing to do with this. 2. The author persist to retain the part of Heterozygosity, but this contribution to conservation is very very insignificant. The authors should at least clarify the limitation for it. 3. The author persist to retain the part of Myostatin. I agree that this is helpful to breeding but one individual could provide very very little solid genetic information for it. The authors should also clarify the limitation.	

GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Editor Decision
Zhang Hongfang Dr Handling Editor
08 8 2024
08 8 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Minor Revision
Nebenführ Marcel Mr Author
15 8 2024
15 8 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Assess Revision
Zhang Hongfang Dr Handling Editor
15 8 2024
16 8 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Final Data Preparation
Tuli Mary-Ann Ms Curator
22 8 2024
09 9 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Editor Decision
Zhang Hongfang Dr Handling Editor
09 9 2024
09 9 2024
GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Accept
Edmunds Scott Dr Editor in Chief
09 9 2024
09 9 2024
Editor’s Assessment	This Data Release paper presents the genome of the whippet breed of dog. Demonstrating a streamlined laboratory and bioinformatics workflows with PacBio HiFi long-read whole-genome sequencing that enables the generation of a high-quality reference genome within one week. The genome study being a collaboration between an academic biodiversity institute and a medical diagnostic company. The presented method of working and workflow providing examples that can be used for a wide range of future human and non-human genome projects. The final is 2.47 Gbp assembly being of high quality - with a contig N50 of 55 Mbp and a scaffold N50 of 65.7 Mbp. This reference being scaffolded into 39 chromosome-length scaffolds and the annotation resulting in 28,383 transcripts. The results also looked at the Myostatin gene which can be used for breeding purposes, as these heterozygous animals can have an advantage in dog races. The reviewers making the authors clarify this part a little better with additional results. Overall this study demonstrating how rapidly animal genome research can be carried out through close and streamlined time management and collaboration.	
Editor’s Assessment	This Data Release paper presents the genome of the whippet breed of dog. Demonstrating a streamlined laboratory and bioinformatics workflows with PacBio HiFi long-read whole-genome sequencing that enables the generation of a high-quality reference genome within one week. The genome study being a collaboration between an academic biodiversity institute and a medical diagnostic company. The presented method of working and workflow providing examples that can be used for a wide range of future human and non-human genome projects. The final is 2.47 Gbp assembly being of high quality - with a contig N50 of 55 Mbp and a scaffold N50 of 65.7 Mbp. This reference being scaffolded into 39 chromosome-length scaffolds and the annotation resulting in 28,383 transcripts. The results also looked at the Myostatin gene which can be used for breeding purposes, as these heterozygous animals can have an advantage in dog races. The reviewers making the authors clarify this part a little better with additional results. Overall this study demonstrating how rapidly animal genome research can be carried out through close and streamlined time management and collaboration.	

GigaByte
GigaByte
Gigabyte
GigaByte
2709-4715
Gigascience Press

Export to Production
Edmunds Scott Dr Editor in Chief
09 9 2024
09 9 2024
==== Refs
References

1 Formenti G , Theissinger K , Fernandes C The era of reference genomes in conservation genomics. Trends Ecol. Evol., 2022; 37 : 197–202. doi:10.1016/j.tree.2021.11.008.35086739
2 Theissinger K , Fernandes C , Formenti G How genomics can help biodiversity conservation. Trends Genet., 2023; 39 : 545–559. doi:10.1016/j.tig.2023.01.005.36801111
3 Guhlin J , Le Lec MF , Wold J Species-wide genomics of kākāpō provides tools to accelerate recovery. Nat. Ecol. Evol., 2023; 7 : 1693–1705. doi:10.1038/s41559-023-02165-y.37640765
4 Mao Y , Harvey WT , Porubsky D Structurally divergent and recurrently mutated regions of primate genomes. Cell, 2024; 187 : 1547–1562.e13. doi:10.1016/j.cell.2024.01.052.38428424
5 Gao H , Hamp T , Ede J The landscape of tolerated genetic variation in humans and primates. Science, 2023; 380 : eabn8153. doi:10.1126/science.abn8197.
6 Cheng H , Concepcion GT , Feng X Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods, 2021; 18 : 170–175. doi:10.1038/s41592-020-01056-5.33526886
7 Astashyn A , Tvedte ES , Sweeney D Rapid and sensitive detection of genome contamination at scale with FCS-GX. Genome Biol., 2024; 25 : 60. doi:10.1186/s13059-024-03198-7.38409096
8 Chen Y , Zhang Y , Wang AY Accurate long-read de novo assembly evaluation with Inspector. Genome Biol., 2021; 22 : 312. doi:10.1186/s13059-021-02527-4.34775997
9 Kolmogorov M , Yuan J , Lin Y Assembly of long, error-prone reads using repeat graphs. Nat. Biotechnol., 2019; 37 : 540–546. doi:10.1038/s41587-019-0072-8.30936562
10 Alonge M , Lebeigle L , Kirsche M Automated assembly scaffolding using RagTag elevates a new tomato system for high-throughput genome editing. Genome Biol., 2022; 23 : 258. doi:10.1186/s13059-022-02823-7.36522651
11 Xu M , Guo L , Gu S TGS-GapCloser: A fast and accurate gap closer for large genomes with low coverage of error-prone long reads. GigaScience, 2020; 9 : giaa094. doi:10.1093/gigascience/giaa094.32893860
12 Vaser R , Sović I , Nagarajan N Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res., 2017; 27 : 737–746. doi:10.1101/gr.214270.116.28100585
13 Rhie A , Walenz BP , Koren S Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol., 2020; 21 : 245. doi:10.1186/s13059-020-02134-9.32928274
14 Simão FA , Waterhouse RM , Ioannidis P BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics, 2015; 31 : 3210–3212. doi:10.1093/bioinformatics/btv351.26059717
15 Li H . Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 2018; 34 : 3094–3100. doi:10.1093/bioinformatics/bty191.29750242
16 Danecek P , Bonfield JK , Liddle J Twelve years of SAMtools and BCFtools. GigaScience, 2021; 10 : giab008. doi:10.1093/gigascience/giab008.33590861
17 Tarasov A , Vilella AJ , Cuppen E Sambamba: fast processing of NGS alignment formats. Bioinformatics, 2015; 31 : 2032–2034. doi:10.1093/bioinformatics/btv098.25697820
18 García-Alcalde F , Okonechnikov K , Carbonell J Qualimap: evaluating next-generation sequencing alignment data. Bioinformatics, 2012; 28 : 2678–2679. doi:10.1093/bioinformatics/bts503.22914218
19 Korneliussen TS , Albrechtsen A , Nielsen R . ANGSD: Analysis of Next Generation Sequencing Data. BMC Bioinform., 2014; 15 : 356. doi:10.1186/s12859-014-0356-4.
20 Smit AFA , Hubley R . RepeatModeler Open-1.0. 2008–2015; http://www.repeatmasker.org/RepeatModeler/.
21 Smit A , Hubley R , Green P . RepeatMasker. Open-4.0. 2013–2015; https://www.repeatmasker.org/.
22 Keilwagen J , Hartung F , Grau J . GeMoMa: homology-based gene prediction utilizing intron position conservation and RNA-seq data. Methods Mol. Biol., 2019; 1962 : 161–177. doi:10.1007/978-1-4939-9173-0_9.31020559
23 Mosher DS , Quignon P , Bustamante CD A mutation in the myostatin gene increases muscle mass and enhances racing performance in heterozygote dogs. PLOS Genet., 2007; 3 : e79. doi:10.1371/journal.pgen.0030079.17530926
24 Steinegger M , Söding J . MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol., 2017; 35 : 1026–1028. doi:10.1038/nbt.3988.29035372
25 Robinson JT , Thorvaldsdóttir H , Winckler W Integrative genomics viewer. Nat. Biotechnol., 2011; 29 : 24–26. doi:10.1038/nbt.1754.21221095
26 Nebenführ M , Prochotta D , Ben Hamadou A Supporting data for “High-speed whole-genome sequencing of a Whippet: Rapid chromosome-level assembly and annotation of an extremely fast dog’s genome”. GigaScience Database, 2024; 10.5524/102573.
