
==== Front
Ecol Evol
Ecol Evol
10.1002/(ISSN)2045-7758
ECE3
Ecology and Evolution
2045-7758
John Wiley and Sons Inc. Hoboken

10.1002/ece3.70302
ECE370302
ECE-2024-07-01412.R1
Microbial Ecology
Microbiomics
Research Article
Research Article
Sampling fish gut microbiota ‐ A genome‐resolved metagenomic approach
Thormar et al.
Thormar Eiríkur A. https://orcid.org/0000-0003-3897-5981
1 eirikur.andri@sund.ku.dk

Hansen Søren B. 1
Jørgensen Louise von Gersdorff 2
Limborg Morten T. 1
1 Globe Institute, Faculty of Health and Medical Sciences, Center for Evolutionary Hologenomics University of Copenhagen Copenhagen K Denmark
2 Section for Parasitology and Aquatic Pathobiology, Department of Veterinary and Animal Sciences, Faculty of Health and Medical Sciences University of Copenhagen Frederiksberg C Denmark
* Correspondence
Eiríkur A. Thormar, Globe Institute, Faculty of Health and Medical Sciences, Center for Evolutionary Hologenomics, University of Copenhagen, Copenhagen K 1353, Denmark.
Email: eirikur.andri@sund.ku.dk

17 9 2024
9 2024
14 9 10.1002/ece3.v14.9 e7030215 8 2024
11 7 2024
29 8 2024
© 2024 The Author(s). Ecology and Evolution published by John Wiley & Sons Ltd.
https://creativecommons.org/licenses/by/4.0/ This is an open access article under the terms of the http://creativecommons.org/licenses/by/4.0/ License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited.

Abstract

Despite a surge in microbiota‐focused studies in teleosts, few have reported functional data on whole metagenomes as it has proven difficult to extract high biomass microbial DNA from fish intestinal samples. The zebrafish is a promising model organism in functional microbiota research, yet studies on the functional landscape of the zebrafish gut microbiota through shotgun based metagenomics remain scarce. Thus, a consensus on an appropriate sampling method accurately representing the zebrafish gut microbiota, or any fish species is lacking. Addressing this, we systematically tested four methods of sampling the zebrafish gut microbiota: collection of faeces from the tank, the whole gut, intestinal content, and the application of ventral pressure to facilitate extrusion of gut material. Additionally, we included water samples as an environmental control to address the potential influence of the environmental microbiota on each sample type. To compare these sampling methods, we employed a combination of genome‐resolved metagenomics and 16S metabarcoding techniques. We observed differences among sample types on all levels including sampling, bioinformatic processing, metagenome co‐assemblies, generation of metagenome‐assembled genomes (MAGs), functional potential, MAG coverage, and population level microdiversity. Comparison to the environmental control highlighted the potential impact of the environmental contamination on data interpretation. While all sample types tested are informative about the zebrafish gut microbiota, the results show that optimal sample type for studying fish microbiomes depends on the specific objectives of the study, and here we provide a guide on what factors to consider for designing functional metagenome‐based studies on teleost microbiomes.

Despite a surge in microbiota‐focused studies in teleosts, few have reported functional data on whole metagenomes as it has proven difficult to extract high biomass microbial DNA from fish intestinal samples. Addressing this, we systematically tested four methods of sampling the zebrafish gut microbiota—collection of faeces from the tank, the whole gut, intestinal content, and the application of ventral pressure to facilitate extrusion of gut material—along with water samples as an environmental control to address the potential influence of the environmental microbiota on each sample type. Using genome‐resolved metagenomics and 16S metabarcoding, our analyses indicate that while all sample types tested are informative about the zebrafish gut microbiota, the optimal sample type for studying fish microbiomes depends on the specific objectives of the study, and here we provide a guide on what factors to consider for designing functional metagenome‐based studies on teleost microbiomes.

fish metagenomes
genome‐resolved metagenomics
sampling fish microbiomes
shotgun metagenomic sequencing
zebrafish gut microbiota
Danmarks Grundforskningsfond 10.13039/501100001732 CEH ‐ DNRF143 Carlsbergfondet 10.13039/501100002808 CF21‐0356 source-schema-version-number2.0
cover-dateSeptember 2024
details-of-publishers-convertorConverter:WILEY_ML3GV2_TO_JATSPMC version:6.4.8 mode:remove_FC converted:17.09.2024
Thormar, E. A. , Hansen, S. B. , Jørgensen, L. v. G. , & Limborg, M. T. (2024). Sampling fish gut microbiota ‐ A genome‐resolved metagenomic approach. Ecology and Evolution, 14 , e70302. 10.1002/ece3.70302
==== Body
pmc1 INTRODUCTION

The application of microbiome data has seen an immense growth in the field of molecular ecology including numerous fish species (François‐Étienne et al., 2023; Minich et al., 2022; Stagaman et al., 2020). However, compared with warm‐blooded animals, there is a remarkable scarcity of microbiome studies on fish that are based on functional insights from whole genome shotgun sequencing techniques. Interestingly, this seems to be partly explained by a significantly lower microbial biomass in fish intestines compared with warm‐blooded animals (Limborg et al., 2023), which has complicated the sampling of microbial DNA as samples often contain >90% host DNA. Here, we argue that part of this challenge can be addressed with an optimised sampling protocol by showcasing the outcome of different protocols to sample a fish gut microbiome community using the zebrafish as a model system.

Zebrafish (Danio rerio) is one of the most studied model organisms to date. The use of the model dates back to the research of George Streisinger in the late 1960s and is today used in a wide variety of research topics from evolutionary biology to human health studies (Grunwald & Eisen, 2002). Today's extensive use of the model is well grounded in that zebrafish are relatively easy to rear, cheap to care for, easy to breed, have a generation time of only 3 months, have a well‐defined embryology (Spence et al., 2008), and are readily amenable to genetic modification (Hwang et al., 2013; Tennant et al., 2019). Furthermore, a plethora of zebrafish‐related scientific literature exists with over 53,000 publications to date (PubMed Search Results 2023). It is no surprise then that interest in applying zebrafish as a model organism to study the increasingly popular topic of host–microbiota interactions is on the rise, and a multitude of research papers, preprints, and reviews have been published on the matter, ranging from probiotic effect on behaviour (Davis et al., 2016; Stagaman et al., 2020) to effects of the immune cell development gene, irf8, on the zebrafish microbiota (Earley et al., 2018). Expectedly, it seems that the most prevalent method of studying the zebrafish gut microbiota is using 16S metabarcoding, a valuable and powerful tool to investigate and describe the composition, diversity, and taxonomic landscape of the gut microbiota. However, given the increasing interest in the zebrafish model for studying host–microbiota interactions and advances in the development of sequencing technologies and bioinformatic tools, it is interesting to consider the number of studies applying more functional approaches. To date, three publications (Kayani et al., 2021, 2022; Zhang et al., 2021) and a single preprint (Gaulke et al., 2020) appear to form the entirety of the zebrafish short‐read‐based metagenomics literature. Only one of which (Kayani et al., 2022) uses a genome‐resolved metagenomics approach, that is de novo assembly of whole bacterial genomes often termed metagenome‐assembled genomes (MAGs). One of the studies relies on sampling whole intestines (Zhang et al., 2021), while the others all rely on faecal matter sampled directly from the tank to characterise the gut metagenome. However, when faecal matter is collected, an issue with the water microbiota arises. As the sample is acquired directly from the tank, the boundary between the host microbiota and the water microbiota may be blurred, potentially rendering functional and taxonomic insights less than optimal. Although studies on larger fish species using metagenomic shotgun sequencing have included intestinal content and gut mucosal scrapes from various gut regions (Cheaib et al., 2021; Collins et al., 2021; Rasmussen et al., 2023; Riiser et al., 2019), those strategies are less feasible for smaller fish, such as zebrafish, which motivates us to focus on sampling strategies with broader applicability across the teleost species.

This leads to considerations regarding which sampling strategy is best for representing the zebrafish intestinal microbiota for a metagenomic shotgun sequencing study while providing enough microbial DNA reads. There are several different methods noted in the literature for sampling the zebrafish microbiota, (Stagaman et al., 2020) but little tangible consensus on what constitutes an appropriate sampling protocol. We therefore addressed the sampling of the zebrafish intestinal microbiota by comparing four different methods throughout a genome‐resolved metagenomic pipeline. We compared faeces sampled directly from the tank (here “faeces”), whole gastrointestinal tract samples (here “whole gut”), pressure on the underside (here ‘squeezed gut’), and finally intestinal content dissected from whole guts (here ‘intestinal content’). Water samples were included as an environmental control and for estimation of environmental contamination. Differences among sample types were evident on all levels from sampling to analyses. There are more potential sampling techniques that may cover the niches of the zebrafish gut microbiota, that is, differential composition and function in different sections of the gut. We have only addressed the aforementioned sampling techniques with an aim on how to best generate whole metagenome data as they are both easy to apply and representative of the whole zebrafish gut microbiota community.

2 METHODS

2.1 Zebrafish husbandry and ethics statement

The zebrafish were reared in a re‐circulatory system (Techniplast active blue) with a light/dark cycle of 14/10 h. The zebrafish were fed with dry feed (ZM‐300, ZM‐400, ZM Fish food and Equipment, UK) and live Brine Shrimp (Artemia sp., ZM Fish food and Equipment, UK). The zebrafish collected for this study were all housed in the same tank. All gut microbiota samples were obtained from adult zebrafish The experiment was conducted in accordance with a permit from The Animal Experiments Inspectorate under the Danish Ministry of Environment and Food (Permit: 2021‐15‐0201‐00951). For euthanisation of fish, 300 mg/L of tricaine methanesulphonate (MS222, A5040, Sigma‐Aldrich, Denmark) was used.

2.2 Sampling

As described above, we aimed to compare four different approaches for their efficiency in sampling DNA for characterising the zebrafish gut metagenome using shotgun sequencing. These sampling methods all differ in degree of invasiveness and relative ease of sampling (Table 1).

TABLE 1 Overview of the sampling process among sample types in terms of feasibility of sampling.

	Dissected whole gut

	Dissected intestinal content

	Squeezed gut content

	Faeces

	
Sample description

	Whole gut dissected from fish	Whole gut dissected from fish and intestinal content dissected out	Pressure applied to the abdominal side of the fish from the pectoral fin to the anal fin and the excretion sampled	Faecal matter collected from the bottom of the tank	
Invasiveness	High	High	Moderate	None	
Ease of sampling	Moderate	Difficult	Moderate	Easy	

2.2.1 Whole gut

The whole gastrointestinal tract of zebrafish was carefully dissected out and stored in an MP Biomedicals Lysing E Matrix tube with 0.5 mL of 1X DNA/RNA Shield (Zymo research). A total of three whole gastrointestinal tracts were collected. This is an invasive, and rather time‐consuming technique but does carry the benefit of individual association, consistency, and no direct contact with the surrounding water. The main drawback of this method is that most of the DNA extracted will likely be host DNA, a common feature of fish microbiota studies (Collins et al., 2021; Hennersdorf et al., 2016; Rasmussen et al., 2021; Riiser et al., 2019), and something we have previously observed in whole guts of zebrafish (unpublished data). This is a potential limitation as high host DNA fractions can impact the quality and resolution of the analysis of the microbiota (Pereira‐Marques et al., 2019).

2.2.2 Squeezed gut

We sampled intestinal content from the zebrafish by gently ‘squeezing’/putting pressure along the abdominal side of the fish from the pectoral fin toward the anal fin. We sampled the resulting excretion from the urogenital opening using an inoculation loop, ~ ½ of a 10 μL loop faecal matter was then transferred to and stored in an MP Biomedicals Lysing E Matrix tube with 0.5 mL of 1× DNA/RNA Shield (Zymo research). A total of three ‘squeezed gut’ samples were collected. Although moderately easy to sample, the samples were inconsistent where a single sample clearly yielded more ‘clean’ faecal matter than did the others.

2.2.3 Intestinal content

The gastrointestinal tract was carefully dissected out and the tip of a single‐use inoculation loop was gently traced along the entire dissected gastrointestinal tract to extrude the intestinal contents avoiding tearing of the intestinal tract, as similarly done in (Gaulke et al., 2016). Using an inoculation loop, ~ ½ of a 10 μL loop of the intestinal content was then transferred to an MP Biomedicals Lysing E Matrix tube with 0.5 mL of 1× DNA/RNA Shield (Zymo research). Four intestinal content samples were collected. Although this method is rather time‐consuming and invasive it has the potential to get both individual‐level variation and potentially more microbial DNA by avoiding host tissue. The four samples were consistent apart from a single sample which yielded substantially more material compared with the others.

2.2.4 Faeces

Faecal matter was sampled directly from the tank using a single‐use polyethylene pipette and transferred to a Petri dish. Using an inoculation loop, ~ ½ of a 10 μL loop, faecal matter was then transferred to and stored in an MP Biomedicals Lysing E Matrix tube with 0.5 mL of 1× DNA/RNA Shield (Zymo research). A total of three replicate faecal samples were collected. This is likely the most prevalent method as it is the easiest and most consistent sample to acquire and completely noninvasive; this sampling method even allows for a time series study to be conducted. However, this approach provides a tank‐level representation of the zebrafish microbiota, lacking individual fish representation. This problem can however be overcome as done by (Gaulke et al., 2019) where individual fish are separated into smaller tanks and faecal matter is subsequently collected.

2.2.5 Water

As environmental control, water samples (n = 3) were collected. For each water sample, 500 mL water was collected, 50 mL at a time and filtered using a sterivex filter. Each sterivex filter was then filled with DNA/RNA Shield (Zymo research) before storage. The water samples were stored at −20°C before DNA extraction.

2.3 DNA extraction

Prior to DNA extraction, all samples were lysed in a Tissuelyzer (Qiagen) at 30 GHz for 5 min and spun down at 16,000g for 5 min. Subsequently, 400 μL lysate from each sample was transferred to a 96 deep well plate in a randomised order. Three DNA/RNA shield (Zymo research) negative extraction controls were included in the process. DNA was extracted following the manufacturer's recommendations of the magnetic bead‐based Quick‐DNA MagBead Plus Kit (Zymo Research). The DNA concentration of each sample was measured using a Qubit fluorometer (Invitrogen).

2.4 Shotgun metagenomics library preparation and sequencing

Twenty five microliters of the eluted DNA was shipped to Novogene (Cambridge, UK) for metagenomic shotgun sequencing where DNA was randomly sheared into short fragments by sonication. Library was prepared using Novogene NGS DNA Library Prep Set (Cat No. PT004). The library was quantified with a Qubit fluorometer and real‐time PCR and bioanalyzer 2100 (Agilent) was used for size distribution detection. The quantified libraries were then pooled and subjected to 150 bp paired‐end sequencing on a NovaSeq 6000 instrument (Illumina).

2.5 16S amplicon library preparation and sequencing

The remaining 20 μL of DNA was used for generating paired 16S data for all samples. First the V3‐V4 hypervariable region of the 16S rRNA gene was amplified using 341F (5′‐CCTAYGGGRBGCASCAG‐3′) and 806R (5′‐GGACTACNNGGGTATCTAAT‐3) (Yu et al., 2005) with the addition of 20 different tags to the 5′ end of each primer. A negative PCR control was also included. Each reaction was 25 μL and consisted of 13.5 μL ddH20, 2.5 μL 10× TaqGold buffer (GeneAmp®), 2.5 μL Mgcl2 (25 mM), 1.5 μL BSA (20 ng/μL), 0.5 μL dNTPs(10 mM), 0.5 μL DNA polymerase AmpliTaq Gold® (5 U/μL), 2 μL primer mix and 2 μL DNA extract (2 μL ddH20 for the PCR control). The PCR conditions were as follows: Initial denaturation at 95°C for 10 min followed by 30 cycles of denaturation at 95°C for 15 s, annealing at 53°C for 20 s and extension at 72°C for 40 s followed by a final extension at 72°C for 10 min. The amplified PCR products were then pooled in an equimolar fashion based on agarose gel band brightness and purified using the HighPrep PCR Clean‐up System® (MagBio Genomics Inc). Subsequently, the PCR‐free single‐tube metabarcoding library preparation protocol, Tagsteady (Carøe & Bohmann, 2020) was used to construct the library. The library was quantified using the NEBNext® Library Quant Kit for Illumina® (New England Biolabs) by mixing 2 μL 1:10,000 diluted library with 8 μL Quant mastermix with added primers (New England Biolabs). Due to low initial concentration of adapter‐ligated products, we made triplicates of the library, pooled the triplicates, and increased the concentration using a vacuum concentrator, Concentrator plus® (Eppendorf). The library was then sequenced at the GeoGenetics Sequencing Core, University of Copenhagen, Globe institute using an Illumina MiSeq platform, reagent kit v3 at 600 cycles. Negative extraction and PCR controls were included in the library and sequenced.

2.6 Bioinformatic processing of shotgun sequencing reads

The quality of raw reads was assessed using FastQC (Andrews, 2010) and MultiQC (Ewels et al., 2016) for subsequent quality‐filtering steps. AdapterRemoval/v.2.3.3 (Schubert et al., 2016) was used for the removal of low quality reads and adapters. Duplicates were removed using the rmdup function in Seqkit/v.2.3.1 (Shen et al., 2016) and reads re‐paired using bbmap/v.38.84 (Bushnell, 2014). Host reads were filtered out by mapping the quality‐filtered reads to the zebrafish genome (GCA_000002035.4) and the potential contamination was removed by mapping to the human genome (hg38) (Schneider et al., 2017) with minimap2/2.24 (Li, 2018, 2021) using default parameters for short accurate genomic reads. All unmapped reads were kept. The mapped reads were kept in separate files to assess the depth of coverage using the samtools‐depth function (Danecek et al., 2021). The metagenomic reads were co‐assembled on a sample type basis (that is, 5 co‐assemblies) using MEGAHIT/v.1.2.9 (Li et al., 2015) with metagenomic sensitive presets and a minimum contig length of 1000 bp. Quality of assemblies was assessed using Quast/v.5.0.2 (Mikheenko et al., 2018). Further analyses of assemblies and visualisation were performed using the anvi'o platform (Eren et al., 2015, 2021).

For each of the five co‐assemblies, the following was done: (i) anvi'o was used to identify Open Reading Frames (ORFs) using Prodigal/v.2.6.3 (Hyatt et al., 2010). (ii) HMMER/v.3.3 (Finn et al., 2011) was used to identify sets of Singe‐Copy‐core Genes (SCGs) of protista, archaeal, and bacterial origin (Lee, 2019). Hidden Markov Models of SCGs were used for estimating the number of recoverable bacterial genomes in the assemblies and for completion and redundancy estimates of MAGs. (iii) The ORFs were annotated in the anvi'o platform using functions from NCBI's Clusters of Orthologous Groups (COGs) (Tatusov et al., 2003), the KOfam HMM database of KEGG orthologs (KOs) (Aramaki et al., 2020; Kanehisa & Goto, 2000) and Pfams (El‐Gebali et al., 2019). (iv) Kaiju (Menzel et al., 2016) was used to infer the taxonomy of genes with NCBI's non‐redundant protein database ´nr´ and import into the anvi'o framework as described here (https://merenlab.org/2016/06/18/importing‐taxonomy/). (v) For the prediction of the number of genomes in the assemblies, we used the ‘anvi‐display‐contigs‐stats’ function. (vi) Reads were then mapped to contigs using minimap2/2.24 (Li, 2018, 2021) and samtools (Danecek et al., 2021) and stored as BAM files. Each BAM file was profiled in anvi'o sung ‘anvi‐profile’ for estimation of coverage and detection statistics of each contig, and each profile was integrated into a merged profile database for further processing using ‘anvi‐merge’. (vii) Finally, Binning was performed using the anvi‐cluster‐contigs which uses CONCOCT/v.1.1.0 (Alneberg et al., 2014). Each bin was then manually refined based on tetranucleotide frequency and differential coverage across samples using ‘anvi‐refine’. MAGs were called using ‘anvi‐rename‐bins’ where each bin that was more than 50% complete and less than 10% redundant was defined as a MAG. Anvi'o was also used to infer MAG taxonomy based on single‐copy core gene sequences from the Genome Taxonomy Database (GTDB) (Parks et al., 2018).

To create a non‐redundant set of MAGs representative of all samples, sequences for each MAG from each merged profile database were extracted. CheckM (Parks et al., 2015) was used as an additional quality assessment of the MAGs. We then applied dRep (Olm et al., 2017) to dereplicate the set of MAGs into a set of non‐redundant species‐level representatives of the collection based on average nucleotide identity (ANI) and quality of MAGs based on the previous CheckM quality assessment (Data S1). We then mapped the metagenomic reads from all samples to the set of non‐redundant MAGs, and followed the same steps as detailed above using the anvi'o platform. For further comparison with existing data, we mapped our metagenomic reads to the Zebrafish Fecal v1.0 MAG catalogue in the MGnify genomes resource (Gurbich et al., 2023).

2.7 Bioinformatic processing of 16S amplicon sequencing data

Quality of raw reads from 16S V3–V4 sequencing was assessed using FastQC (Andrews, 2010) and MultiQC (Ewels et al., 2016) for subsequent quality filtering steps. Cutadapt/v.4.4 (Martin, 2011) was used to trim adapter and primer sequences. Subsequently, the DADA2/v1.26 (Callahan et al., 2016) pipeline was applied for quality filtering and trimming, inferring of amplicon sequence variants (ASVs), merging of forward and reverse reads, removal of chimeric sequences, and taxonomy assignment with the SILVA/v138 reference database (Quast et al., 2013). Then Decontam/v1.20.0 (Davis et al., 2018) was applied to remove putative contaminants, after which LULU/v0.1.0 (Frøslev et al., 2017) was applied for ASV curation and removal of erroneous sequences. The ASVs were then merged at the genus level for analyses.

2.8 Analyses and visualisation

Rarefaction curves were estimated using the R package vegan (Oksanen et al., 2020). The relative abundances of MAGs were calculated using the mean coverage across the MAG and then normalised using the length of the MAG using TPM normalisation with the R package ADI‐impute (Xu et al., 2021). Microbiota composition analyses were performed using the R package phyloseq (McMurdie & Holmes, 2013). Visualisation of data was performed using the R package ggplot2 (Wickham, 2016) the anvi'o platform (Eren et al., 2015, 2021). The icons representing sample types, used in figures, were created using BioRender.com.

3 RESULTS

3.1 Sample types differed in all stages of sample processing

The sample types differed at every stage of sample processing in data output and quality. Here we refer to sample processing as the process of generating MAGs from DNA extraction through bioinformatic processing and metagenomic assembly. After comparing the practicality of sampling (Table 1), we compared the average post‐extraction DNA concentration among the different sample types (Figure 1a). As expected, due to the high host DNA content, the whole gut had the highest average DNA concentration, followed by intestinal content and squeezed gut, while faeces and water had the lowest average concentration.

FIGURE 1 Comparison among sample types through the sample processing. The colours and icons represent different sample types. From top to bottom, they represent whole gut, intestinal content, squeezed gut, faeces, and water. (a) Mean DNA concentration among sample types. (b) Mean percentage of reads removed after quality filtering and deduplication. (c) Mean percentage of quality filtered and deduplicated reads removed after host filtering. (d) Mean depth of coverage of the zebrafish reference genome. (e) Number of contigs in co‐assemblies among sample types (f) N50 value of the co‐assemblies among sample types, and (g) the GC content of the co‐assemblies among sample types.

Metagenomic shotgun sequencing yielded ~ 2.4 billion reads across 15 samples. A single water sample failed sequencing due to low DNA content along with a negative control (indicating the efficacy of the negative control) and these two samples were therefore not included in the analysis. After quality filtering and removal of duplicates ~ 647.5 million reads remained. The average percentage of reads removed after quality filtering and removal of duplicates was then compared among sample types (Figure 1b). Whole gut samples had the lowest percentage of removed reads followed by intestinal content, squeezed gut, water, and finally most reads were removed from the faecal samples. The standard deviations should also give an indication of the consistency within the sample type group making the intestinal content and squeezed gut less consistent than whole gut and faecal samples. After human and zebrafish filtering ~ 157 million reads remained. The percentage of quality‐filtered reads that were removed by host filtering was then compared among the sample types to give an indication of the amount of host content among sample types (Figure 1c). As expected, the highest amount removed by the host filter was whole gut samples with a strikingly high percentage (99%), followed by intestinal content (86.7%), squeezed gut (84.5%), water (1.29%), and interestingly, the lowest percentage was removed from the faeces samples (0.473%). Notably, the sheer number of reads not mapping to the host zebrafish genome was highest in the water samples followed by faeces, intestinal content, squeezed gut, and finally whole gut. The data therefore counterintuitively indicate that among the sampling types, a lower initial post‐extraction DNA concentration yields more metagenomic reads in the case of zebrafish shotgun metagenomic sequencing.

Despite this, recovering host genetic data may also be a valuable resource for some studies and for avoiding waste of sequencing depth and cost. We therefore included a comparison of the average depth of coverage of the zebrafish genome recovered from each sample type (Figure 1d). Surprisingly, faecal samples presented an unexpectedly lower coverage compared with the water samples. Squeezed gut and intestinal content exhibited relatively similar depths of coverage. As anticipated, the whole gut samples had the highest depth of coverage.

Metagenomic assemblies of the unmapped reads were made for all sample types and compared (Figure 1e–g). The assemblies differed at all levels, the number of contigs was by far the highest in the water samples followed by faeces, squeezed gut, intestinal content, and finally, the lowest number of contigs was observed in the assembly of the whole gut. As depicted (Figure 1f,g), the GC content and N50 value varied widely among the assemblies of the different sample types.

3.2 Faecal samples provide the highest number and quality of MAGs among sample types

Based on the presence/absence of 71 SCGs, the estimated total number of bacterial genomes in the five assemblies was estimated (Figure 2a). The lowest estimated number of genomes was as expected in the whole gut samples (2 bacterial genomes) followed by, intestinal content (13 genomes), squeezed gut (25 genomes), faeces (24 genomes), and finally water had the highest number of estimated bacterial genomes (49 genomes). After automated binning and manual refinement of the MAGs, we compared the number and quality of MAGs. MAGs were called if they were over 50% complete and had less than 10% redundancy, MAGs were classified as high quality if they were over 90% complete and less than 5% redundant. Low‐quality MAGs less than 50% complete and higher than 10% redundant were not included in the analysis. In total 77, MAGs were called across sample types. The highest number of MAGs (28 MAGs) were called in the water samples (57% of expected MAGs) of which 14 were classified as high quality. A total of 23 MAGs were called from the faecal samples (95% of expected MAGs) of which 12 were classified as high quality. From the squeezed gut samples, 16 MAGs were called (64% of expected MAGs) 11 of which were classified as high quality. Eight MAGs were called in the intestinal content samples (61% of expected MAGs) of which 5 were classified as high quality. Finally, only two MAGs were called in the whole gut samples, one of which was classified as high quality. It is clear that the representativeness, the number and quality of the assemblies vary widely, so to compare the sample types, we created a MAG catalogue consisting of 41 non‐redundant MAGs (see Data S1) and independently mapped host‐filtered reads from all samples back to it.

FIGURE 2 (a) Plot showing the estimated and actual number of bacterial genomes in the co‐assemblies among sample types. The circles represent the number of called MAGs and the squares represent the estimated number of MAGs. (b) Plots showing high‐quality and lower‐quality MAGs. The triangle represents high‐quality MAGs (>90% complete and <%5 redundant) while the diamond represents lower‐quality MAGs (>50% complete and >10% redundant). The colours and icons indicate the sampling type, yellow represents whole gut samples, red represents Intestinal content samples, purple colour represents squeezed gut samples, brown represents faeces and blue represents water samples.

First, we assessed sequencing depth and generated a rarefaction curve of gene calls across all samples which revealed striking and interesting differences among sample types in terms of data saturation (Figure 3a). Expectedly, as only two MAGs were recovered, the whole gut samples did not reach saturation, indicating insufficient sampling depth. We observed that the intestinal content samples reached saturation indicating sufficient sampling depth. This is in stark contrast to the squeezed gut samples where most of the diversity seems to belong to a single sample while the others do not reach saturation, indicating a lack of consistency in sampling. However, the saturated curve from the single squeezed gut sample is similar to the ones from the faecal samples which all saturate around the same time as do the intestinal content samples although at a much higher diversity. The water samples are almost saturated and indicate the highest diversity. This is supported by the coverage of the MAG catalogue among samples.

FIGURE 3 MAG abundance and diversity among sampling types. The colours and icons indicate the sampling type. (a) Rarefaction curve of ORFs among sample types. (b) PCoA of MAG composition among samples using Bray‐Curtis distance with the colours representing sample type as in A. (c) PCoA of MAGs where the blue coloured gradient represents the percentage of removed host reads. (d) log transformed mean per sample coverage of MAGs in the MAG catalogue created for this study. Each column of squares represents a single MAG and each row represents a single sample. GC content of each mag is represented by black bars and completion and redundancy estimates of each MAG is represented by blue and red bars respectively, the genus level of each MAG is also indicated in text. The three MAGs representing over 5% of the overall abundance are indicated with asterixis on the dendrogram.

We then looked at the relative abundances of taxa across the dataset (Figure 3d). Of the 41 MAGs, only three comprise over 5% of the relative abundance across the sample set, one of which belongs to the genus Cetobacterium (assigned species Cetobacterium sp000799075) is clearly dominant across all sample types. It comprises the highest percentage in the whole gut samples (85.2%) followed by intestinal content (72.1%), squeezed gut (67.4%), faeces (39.8%), and water (38.7%). The second most abundant MAG belongs to the genus Aeromonas and its relative abundance is highest in the faecal samples (13.2%), followed by water (12.4%), intestinal content (11.2%), squeezed gut (5.03%), and whole gut (3.62%). The third and final MAG comprising over 5% of the relative abundance also belongs to the genus Cetobacterium (assigned species Cetobacterium somerae). Its relative abundance is highest in the faecal samples (8.38%), followed by intestinal content (7.36%), water (6.19%), squeezed gut (6.156%), and finally whole gut (4.04%). A total of 10 MAGs (the three summarised included) comprise over 1% of the cumulative relative abundance (a summary of those can be found in Data S1). It should be noted that one of the 10 MAGs has by far the highest relative abundance in the water samples, that is a MAG belonging to the genus Ideonella. Despite the obvious differences among sample types in terms of relative abundance, significant differences were not observed among the sample types, likely due to the small per‐group sample numbers. The beta‐diversity of the microbiota composition shows significant differences based on sample type (Figure 3b). The first axis of the PCoA separates the whole gut, intestinal content, and part of the squeezed gut content from the faecal and water samples and a single squeezed gut sample (indeed the one that looks more like the faecal samples). The second axis of the PCoA separates the water samples reflecting the different composition of the water compared with the other samples. Moreover, there appears to be a gradient in samples ranging from low‐high host DNA content along the first axis of the PCoA (Figure 3c).

Before analysing the functional potential among the sample types we compared the relative abundances of MAGs to 16S amplicon‐based ASVs generated for the same samples (Figure 4) We focused on comparing the relative abundance of four prevalent taxa ensuring a consistent genus‐level match between the 16S dataset and the MAG catalogue. These taxa belong to the genera Cetobacterium, Aeromonas, Pseudomonas, and ZOR0006. Although not highly significant in all samples, the correlation between log2 transformed relative abundances of the four taxa in 16S and shotgun seems almost perfect among sample types indicating that the functional metagenomic data accurately represents the composition of the microbiota as characterised by the normally used 16S‐based barcoding approach. The main difference seems to be in the abundances of Cetobacteriuma and Aeromonas in the faecal and water samples of which there is a slightly higher abundance in the 16S data compared with the shotgun data (see also Supplementary file S1, Figure S1) and in the Pseudomonas which is consistently higher in abundance in the shotgun data. Diversity and composition analyses were also performed on the 16S data and they largely reflected the composition of the shotgun data although there was less difference between the sample types, that is, the 16S data from the faecal samples was similar to the whole gut samples (Supplementary file S1, Figure S1).

FIGURE 4 This figure shows the relationship between the log2 transformed relative abundances of the top 4 taxa shared between the 16S metabarcoding data and the shotgun sequencing data among the five sample types. The line represents an x = y relationship. The icons represent the different sampling types and the colours differentiate the four genera.

3.3 Functional composition differs among sample types

Having established the differences in composition among sample types and validation with 16S metabarcoding data, we aimed to investigate the differences in the functional composition. From the difference between the abundance of MAGs among sample types, we expect the gene content to vary similarly among the sample types with increasing depth of functional inference (Figure 3). We started by comparing COG categories among sample types (Figure 5). A total of 168,366 genes were called across the MAG catalogue and of those 120,410 had assigned COG categories. We compared the sample types in how many gene calls were unique/shared among sample types (Figure 5a) and how much potential functional fraction (that is the fraction of normalised coverage of gene calls) those gene calls contribute (Figure 5b). We identified three main fractions of gene calls among sample types. Firstly, a core fraction of gene calls that are shared among all sample types. Secondly, a fraction shared among all sample types except for whole gut samples was included, likely reflecting the diversity of functions not captured by the whole gut content samples due to high host DNA content. And lastly, a fraction is shared exclusively between faecal samples and water samples. All sample types shared a core fraction of 35,282 gene calls with assigned COG functions. This fraction of gene calls accounts for most of the potential functional abundance 99.7% (SD = 0.0974) in the whole gut samples, 98.8% (SD = 0.359) in the intestinal content samples, 95.2% (SD = 3.49) in the squeezed gut samples, 90.1% (SD = 3.77) in the faecal samples and only 45.1% (SD = 7.58) in the water samples suggesting that these functions likely represent a set of core functions in the zebrafish gut microbiota opposed to be sourced from the water microbial community. In the fraction shared among all samples except the whole gut samples (20,967 gene calls), the fraction of functional potential was only 1.02% (SD = 0.269) in the intestinal content samples, 4.2% (SD = 3.26) in the squeezed gut samples, 6.79% (SD = 1.51) in the faecal samples and 5.16% (SD = 0.0203) in the water samples. Thus, likely reflecting the increased resolution of the microbial functions of lower abundance taxa when avoiding sampling of the host gut tissue. Looking at the fraction, only shared among the water and the faeces (25,736 gene calls), it accounts for only 1.19% (SD = 0.728) of the functional potential in the faecal samples in stark contrast with the 33.2% (SD = 5.34) in the water samples. This likely reflects a fraction which corresponds to contamination from the water samples. We then removed the fraction of gene calls from the sample set that had a higher mean coverage in the water samples compared with the other sample types. This resulted in a big reduction of shared gene calls both in the fractions shared among only the faecal samples and water samples and in the fraction shared among all sample types apart from whole gut samples. Further addressing this, we focused only on the gene calls with COG functions from the 10 MAGs that constituted over 1% of the relative abundance (24,432 gene calls). Subtracting the gene calls from the Ideonella MAG, which had the highest abundance in the water samples completely removed the fraction of gene calls (2604 gene calls) shared between only the faecal and water samples along with a part of the fraction shared between all sample types except the whole gut samples. This analysis was repeated for gene calls with assigned Pfam functions (Supplementary material file S1, Figure S2) and indicated similar results.

FIGURE 5 (a) Venn diagram of the number of gene calls with assigned COG categories across sample types. (b) Percentage of normalised gene coverage among sample types in fractions of the Venn diagram. The respective fractions are indicated by arrows. The colours and icons represent the different sample types. The colours and icons indicate the sampling type, yellow is whole gut samples, red represents the intestinal content samples, purple is squeezed gut samples, brown represents faeces and blue represents water samples.

3.4 Intra‐MAG variability is driven by MAG coverage

To briefly address the increasingly insightful and rising topic of microbial population genetics, we profiled the single‐nucleotide variants (SNVs) variability among sample types (Figure 6). To do this, we chose the highest abundant MAG across the sample set, namely, the Cetobacterium sp000799075 MAG. Only SNVs with more than 10X coverage across 90% of the samples were included. The results indicate that the degree of variability is mainly driven by the coverage of the Cetobacterium sp000799075 MAG (Pearson correlation R = .91, p = 1.9 × 10−6). That is, the higher the MAG coverage (and the lower the host DNA content depending on sample type), the more variation is recovered for each SNVs indicating that too low coverage of MAGs leads to loss of information about the intraspecific microbial population diversity.

FIGURE 6 This figure shows the variability of single‐nucleotide variants (SNVs) in the high abundance Cetobacterium sp000799075 MAG. SNVs with a minimum coverage of 10× across 90% of the samples were taken into account. The variability of SNVs is hierarchically clustered using ward ordination and Euclidean distance based on the similarity of variability. The vertical ordering is clustered based on similarity among samples. The two columns on the right represent a gradient of the MAG's coverage in each sample (yellow) and a gradient of percentage of host reads in each sample.

3.5 The representation of MAG catalogues differ among sample types

We also wanted to investigate how the data generated for this study compared with previously generated zebrafish metagenomic data. We mapped our metagenomic reads to a previously curated MAG catalogue (Gurbich et al., 2023; Richardson et al., 2023) consisting of 101 MAGs. We compared the fraction of metagenomic reads mapping to the Zebrafish Fecal v1.0 MAG catalogue to the percentage mapping to our own MAG catalogue which should estimate both the representativeness of both catalogues with regards to our data and sample types (Figure 7). The overall trend among all sample types was that a higher percentage of reads mapped to the MAG catalogue generated for this study compared with the reference Zebrafish Fecal v1.0 MAG catalogue, with the difference ranging between 5% and 24%. Notably, the highest fraction of reads mapped to both catalogues was observed in the intestinal content samples followed by faecal samples, squeezed gut samples, whole gut samples, and finally the water samples.

FIGURE 7 This figure shows the comparison of unmapped reads generated in this study mapped to the non‐redundant MAG catalogue generated for this study and the Zebrafish Fecal v1.0 MAG catalogue from (Gurbich et al., 2023) which was generated by publicly available zebrafish shotgun metagenomic sequences. Figure created using BioRender.com.

4 DISCUSSION

We aimed to illustrate the impact of sampling technique on bioinformatic processing and analyses of teleost gut microbiota through a typical genome‐resolved metagenomic pipeline. We considered the zebrafish model and compared four sampling techniques as well as water samples serving as an environmental control to assess the potential environmental contamination from water‐related microbes. We compared the sample types on practicality and consistency of sampling, through quality filtering, assembly, MAG recovery, and microbiota composition.

Considering a sample that constitutes a good representative for the zebrafish gut microbiota, sampling the whole gut seemed like a good idea at a first glance since in theory, it should capture the entirety of what is in the gut and even potential intracellular gut bacteria without potential contamination from the environment. However, considering the results, it may not be the best representative sample. The sheer amount of host reads (99%) leads to low number of microbial reads which leads to problems with downstream analyses. Even though we sequenced 20G per sample, the sequencing depth of the microbial content of the whole gut samples was insufficient. We only recovered two MAGs, which is low compared with the rest of the sample types and what we expected from existing literature (Kayani et al., 2021, 2022). Those MAGs, belonging to the genera of Cetobacterium and Aeromonas, were however the ones that dominated the relative abundance across all sample types. This may indicate that despite a low metagenomic content in the whole gut sample, the potential for monitoring the dominating bacteria in the gut is there. We encountered similar problems with sampling the intestinal content, although to a lesser extent.

The content of the zebrafish intestinal tract should, like the whole gut samples, provide a complete representation of the zebrafish gut microbiota. Although more time‐consuming in time spent per sample, the intestinal content had a lower amount of host DNA than the whole gut samples. Although 87% host DNA is a large fraction, it is an improvement from the whole gut samples. Rarefaction analyses indicated sufficient sequencing depth and saturated curves although the diversity is less than the faecal samples. Despite the overall higher quality of data from the intestinal content samples compared with the whole gut samples, the microbiota profiles largely reflect a similar composition to the whole gut samples, if the non‐redundant MAG catalogue is followed, with Cetobacterium and Aeromonas dominating the composition. Overall, the intestinal content samples seem to provide a microbiota resolution that is an intermediate between whole gut and faecal samples.

We included squeezed gut samples as we thought they would be a potentially good alternative to faecal samples. That is, sampling the faeces from the fish without dissecting out the gut and without the faeces coming into contact with the surrounding water, thus largely eliminating the potential problem for host and environmental contamination. However, according to the results, unfortunately, the sampling was inconsistent. Analyses of two of the three samples taken indicated insufficient sequencing depth and a similar host content as in the whole gut samples, in fact, they largely reflected the whole gut samples throughout the bioinformatic processing and analyses. However, a single sample only had 55% host content, indicating sufficient sequencing depth and larger diversity compared with the intestinal content samples. This sample is likely responsible for the majority of the 16 MAGs recovered from the binning efforts of assembled contigs from the squeezed gut samples. We believe that if this sampling method can be improved in terms of consistency, it lessens the risk associated with potential environmental contamination as in the faeces sampled directly from the tank water. Interestingly, the squeezed gut sample with higher diversity largely reflected the composition and diversity of the faecal samples. That sample may therefore represent the microbial fraction not present in the whole gut and intestinal content samples while avoiding the potential water contamination.

Sampling faeces directly from the tank is the most prevalent method that has been used in the existing literature of shotgun metagenomic sequencing of the zebrafish microbiota (Gaulke et al., 2020; Kayani et al., 2021, 2022). The sampling effort indicated very low host contamination (<1%), sufficient sequencing depth and recovery of 23 MAGs, moreover, the three replicate samples were largely consistent throughout the analyses. The faecal samples seem to reveal a microbial fraction not present in the whole gut or intestinal content samples. In this study, the faecal samples were not from individual fish but rather represent a group level within each tank. Although potentially impacting the comparisons among the sample types, the results should be applicable to individual‐based sampling efforts. For obtaining an individual resolution (Gaulke et al., 2019) provide a solution where individual fish were placed in individual tanks overnight and faecal matter sampled. Whether such a method produces an additional discernible tank effect should be investigated. In terms of the faecal sampling strategy used in this study, there are more methods that can be easily applied depending on resources, time, and purpose, such as collecting the faeces from the water, spinning down the sample and removing the supernatant leaving a pellet that can be processed as a faecal sample as done in (Sieler Jr et al., 2023), Although the initial faecal sampling methodology may thus slightly impact the results in terms of quality, considering the low amount of host contamination and that the end result is a faecal sample without host tissue, the results presented here are suitable for other faecal sampling strategies The faecal samples analysed here raised two questions, one regarding the composition of the microbiota, as it was discernibly different from the other samples and secondly regarding environmental contamination.

Starting with the overall microbiota composition it is clear that two MAGs dominate the microbiota composition (the Cetobacterium sp000799075 MAG and Aeromonas MAG), These have been found in high abundance in both previous shotgun metagenomic studies and in 16S‐based studies. This study may simply reflect the microbiota of the zebrafish in the facilities in which the study was conducted, as facility/study‐specific microbiota are well documented (Roeselers et al., 2011; Sharpton et al., 2021). What is more interesting is the difference in composition and diversity between the intestinal content samples and the faecal samples. This might simply reflect the added microbial fraction not captured by the intestinal content samples. It could also be that the faecal samples include microbes that are simply passing through. Otherwise, it is a possibility that the faecal samples and the intestinal content samples just represent different ‘niches’ or different parts of the zebrafish gut microbiota. The number of MAGs generated in this study is perhaps at first glance rather low compared with the existing zebrafish literature, but considering the high host content and the abundance of the most abundant MAGs, this is perhaps to be expected. However, future studies are expected to have larger sample sizes compared with this pilot, and should improve the binning efforts and yield a higher number of MAGs. On the other hand, a similar low number of MAGs have been reported in other metagenomic studies of fish species like Atlantic cod (Le Doujet et al., 2019) and Atlantic Salmon (Cheaib et al., 2021; Rasmussen et al., 2023), seemingly reflecting a general trend of low functional diversity in teleost gut microbiomes (Limborg et al., 2023).

One of the main concerns and basis of this study was the potential contamination of faecal samples from the surrounding water. We therefore attempted to apply a functional approach to the estimation of the extent of contamination. In terms of functional abundance, the results indicated a low amount of contamination in the faecal samples. The results also indicate that the water samples are more contaminated from the faecal samples than the faecal samples are potentially contaminated by the water samples. This may be the case due to a higher biomass in the faecal samples, potentially overwhelming the water microbiota community. One might also assume that MAGs or genes with higher coverage in the water samples should originate from the water; however, one can never exclude the possibility that a higher coverage MAG from the water lives in a lower abundance niche in the zebrafish gut. Therefore, it is difficult to say whether actual water contamination occurs in the faecal samples, or if it is simply something that is present in low abundance in the faecal samples. Furthermore, it is difficult to say whether increased sample numbers and sequencing efforts would potentiate or reduce this contamination risk. These results are based on coverage of genes and thus any analytical approach using presence/absence of genes or metabolic pathways need to be further scrutinised as a large fraction of gene calls is shared between the environmental water microbiota (Figure 5a) and the intestinal sample types. Surprisingly, more host DNA contamination was present in the water samples than the faecal samples, making the use of water samples as a no‐host control potentially inappropriate. The host content in the water from the tank may potentially originate from shedded mucus or other tissues from the fish. Despite the low potential water contamination of the faecal samples, taken together the results indicate that caution should be taken when interpreting results from faecal matter sampled directly from the tank, and that water samples serve as an important environmental control that can help identify and remove MAGs more likely to be sourced from the water.

We further compared our non‐redundant MAG catalogue with a previously published MAG Zebrafish faecal catalogue (Gurbich et al., 2023) consisting of 101 MAGs. Overall, the MAG catalogue generated in this study performed better compared with the existing reference catalogue. This is not surprising as our study catalogue was generated from our data and should thus suit it better. However, the difference between the percentages of reads mapping back to the catalogues was smaller than expected. This demonstrates the usefulness of using other reference MAG catalogues, as having a comprehensive MAG catalogue can make data processing and analyses more straightforward. One may for example capture a higher fraction of reads for downstream analyses otherwise lost in the local assembly process where rare species may not be well represented in study‐specific de novo MAG catalogues. This is particularly relevant to zebrafish considering the facility and study‐related differences documented in the 16S zebrafish microbiota literature (Roeselers et al., 2011; Sharpton et al., 2021).

The number of MAGs recovered per sample type is relatively low compared with existing literature in zebrafish (Gaulke et al., 2020; Gurbich et al., 2023; Kayani et al., 2021, 2022). However, considering the limited per‐group sample number it is expected that the number of MAGs will increase with increased sample size due to increased binning power. The small per group sample sizes of two to four and only 15 samples in total, led us to largely compare the sample types qualitatively. We decided on a small‐sample‐size‐deeper‐sequencing instead of larger‐sample‐size‐shallower‐sequencing approach since we expected a high host DNA content in most of the sample types, based on previously published literature of fish intestinal microbiota studies (Collins et al., 2021; Hennersdorf et al., 2016; Rasmussen et al., 2021; Riiser et al., 2019).

The amount of host DNA is clearly different among the sample types, resulting in lower resolution of the microbial diversity. However, this is not necessarily detrimental. If there is combined interest in the host genome and microbiota, sequencing the whole gut or the squeezed gut may be a good option to retrieve whole host genomes along with the microbial fraction in one sequencing run (Marcos et al., 2022).

In this study, all samples were collected from adult zebrafish. Thus, the study does not cover sampling methodology for shotgun metagenomic sequencing in other life‐stages such as larval zebrafish. Obtaining metagenomic data for the larval fish is a different challenge entirely; therefore, it is unlikely that any of the sampling strategies discussed here would be sufficiently applicable to larval zebrafish, except for whole intestinal samples which can be dissected from the larvae (Rawls et al., 2004; Stagaman et al., 2020). However, this would then likely, similarly to the results presented here, result in a high host DNA content, and low recovery of microbial reads. While larval zebrafish, with their unique sampling challenges, offer invaluable insights into early development of the microbiota, we could see an expansion of focus on the adult zebrafish microbiota as a stable and complementary counterpart in understanding the microbial community dynamics. Although other life‐stages were not addressed in this study, the results presented are likely applicable to juvenile and adult zebrafish along with other small fish species or, for example, amphibians with similar exothermic lifestyles.

For studies simply aiming to study the zebrafish microbiota composition and function, the amount of host DNA in the host‐derived samples is an obvious problem. A simple approach to increase the resolution of the microbiota is to increase the sequencing depth. However, this is poorly scalable, in terms of cost, sample number, data produced, and computational resources. Other methods to reduce host contamination such as chemical host depletion pre or post DNA extraction, microbial enrichment, and host depletion by adaptive sequencing have been applied and compared with varying degrees (Horz et al., 2008; Marotz et al., 2018; Marquet et al., 2022; Yap et al., 2020). These promising alternatives to deep‐sequencing approaches are interesting tools to apply to intestinal zebrafish samples for shotgun or long‐read metagenomic sequencing. As this study simply aimed to compare and evaluate sample types these approaches were not tested. Overall, this study indicates that the choice of sample type matters and underscores the importance of environmental controls, which allows for more scrutiny in analyses. We have attempted to summarise the results from the study (Table 2) with the aim to guide researchers planning to undertake genome‐resolved shotgun metagenomic approaches in zebrafish or other teleost species. We conclude that the method of choice should depend largely upon the aim of the study and that special care has to be taken concerning environmental contamination.

TABLE 2 The table summarises the main results. The colours of the cells correspond to degrees of comparison. Green represents the most favorable ranking, yellow represents the moderate ranking and red represents the least favorable ranking. The light‐red represent the least favourable ranking in the conditions of this study but indicates that there is potential for alternative solutions or optimisation.

	Dissected whole gut

	Dissected intestinal content

	Squeezed gut content

	Faeces

	
Sample description

	Whole gut dissected from fish	Whole gut dissected from fish and intestinal content dissected out	Pressure applied to the abdominal side of the fish from the pectoral fin to the anal fin and the excretion sampled	Faecal matter collected from the bottom of the tank	
Invasiveness

	High	High	Moderate	None	
Ease of sampling	Moderate	Difficult	Moderate	Easy	
Consistency

	High	High	Low a	High	
Sample time

	Long	Long	Moderate	Short	
Individual resolution	Yes	Yes	Yes	No b	
Host dna content	High	High	High a	Low	
Saturated rarefaction curve	No	Yes	Partially a	Yes	
Mag recovery	Poor	Moderate	Moderate	Good	
Water contamination susceptibility	Low	Low	Low	Moderate	
a Potential for optimisation and improvement in sampling strategy.

b Can be used for individual resolution as described by (Gaulke et al., 2019).

AUTHOR CONTRIBUTIONS

Eiríkur A. Thormar: Conceptualization (equal); data curation (equal); formal analysis (equal); investigation (equal); methodology (equal); software (equal); visualization (equal); writing – original draft (lead); writing – review and editing (equal). Søren B. Hansen: Conceptualization (equal); data curation (equal); formal analysis (equal); software (equal); writing – review and editing (equal). Louise von Gersdorff Jørgensen: Conceptualization (supporting); data curation (equal); project administration (equal); resources (equal); supervision (equal); writing – review and editing (equal). Morten T. Limborg: Conceptualization (equal); data curation (equal); funding acquisition (equal); investigation (equal); resources (equal); supervision (equal); writing – original draft (equal); writing – review and editing (equal).

CONFLICT OF INTEREST STATEMENT

The authors have no competing interests to declare.

BENEFIT SHARING STATEMENT

Benefits from this research accrue from the sharing of our data and results on public databases as described above.

Supporting information

Figure S1‐S2.

Data S1.

ACKNOWLEDGEMENTS

This project was funded by The Danish National Research Foundation award (CEH – DNRF143) and a Carlsberg Foundation award to M.T.L. (CF21‐0356). We would like to thank Sara Vebæk Gelskov and Carlota Marola Fernandez Gonzales for assistance with zebrafish husbandry. We would further like to thank Tom O. Delmont at Genoscope & CEH, and Jacob Agerbo Rasmussen for their expert input on the data analyses.

DATA AVAILABILITY STATEMENT

All raw sequences generated for this study in the European Nucleotide Archive (ENA) under the accession number PRJEB71469 (https://www.ebi.ac.uk/ena/browser/view/PRJEB71469) (Thormar, 2023a). An overview of all MAGs generated for the study is publicly available at (https://doi.org/10.6084/m9.figshare.24866772.v1) (Thormar, 2023b). along with fasta files for all MAGs at (https://doi.org/10.6084/m9.figshare.24866790.v1) (Thormar, 2023c), including the non‐redundant set of MAGs at (https://doi.org/10.6084/m9.figshare.24866841.v1) (Thormar et al., 2023). Code used for analysing and processing the data can be found at https://github.com/eirikurandri/Microzebra.
==== Refs
REFERENCES

Alneberg, J. , Bjarnason, B. S. , de Bruijn, I. , Schirmer, M. , Quick, J. , Ijaz, U. Z. , Lahti, L. , Loman, N. J. , Andersson, A. F. , & Quince, C. (2014). Binning metagenomic contigs by coverage and composition. Nature Methods, 11 (11 ), 1144–1146. 10.1038/nmeth.3103 25218180
Andrews, S. (2010). FastQC A quality control tool for high throughput sequence data . https://www.bioinformatics.babraham.ac.uk/projects/fastqc/
Aramaki, T. , Blanc‐Mathieu, R. , Endo, H. , Ohkubo, K. , Kanehisa, M. , Goto, S. , & Ogata, H. (2020). KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold. Bioinformatics, 36 (7 ), 2251–2252. 10.1093/bioinformatics/btz859 31742321
Bushnell, B. (2014). BBMap: A fast, accurate, splice‐aware aligner . https://escholarship.org/uc/item/1h3515gn
Callahan, B. J. , McMurdie, P. J. , Rosen, M. J. , Han, A. W. , Johnson, A. J. A. , & Holmes, S. P. (2016). DADA2: High‐resolution sample inference from Illumina amplicon data. Nature Methods, 13 (7 ), 581–583. 10.1038/nmeth.3869 27214047
Carøe, C. , & Bohmann, K. (2020). Tagsteady: A metabarcoding library preparation protocol to avoid false assignment of sequences to samples. Molecular Ecology Resources, 20 (6 ), 1620–1631. 10.1111/1755-0998.13227 32663358
Cheaib, B. , Yang, P. , Kazlauskaite, R. , Lindsay, E. , Heys, C. , Dwyer, T. , De Noa, M. , Schaal, P. , Sloan, W. , Ijaz, U. Z. , & Llewellyn, M. S. (2021). Genome erosion and evidence for an intracellular niche ‐ exploring the biology of mycoplasmas in Atlantic salmon. Aquaculture, 541 , 736772. 10.1016/j.aquaculture.2021.736772 34471330
Collins, F. W. J. , Walsh, C. J. , Gomez‐Sala, B. , Guijarro‐García, E. , Stokes, D. , Jakobsdóttir, K. B. , Kristjánsson, K. , Burns, F. , Cotter, P. D. , Rea, M. C. , Hill, C. , & Ross, R. P. (2021). The microbiome of deep‐sea fish reveals new microbial species and a sparsity of antibiotic resistance genes. Gut Microbes, 13 (1 ), 1–13. 10.1080/19490976.2021.1921924
Danecek, P. , Bonfield, J. K. , Liddle, J. , Marshall, J. , Ohan, V. , Pollard, M. O. , Whitwham, A. , Keane, T. , McCarthy, S. A. , Davies, R. M. , & Li, H. (2021). Twelve years of SAMtools and BCFtools. GigaScience, 10 (2 ), giab008. 10.1093/gigascience/giab008 33590861
Davis, D. J. , Doerr, H. M. , Grzelak, A. K. , Busi, S. B. , Jasarevic, E. , Ericsson, A. C. , & Bryda, E. C. (2016). Lactobacillus plantarum attenuates anxiety‐related behavior and protects against stress‐induced dysbiosis in adult zebrafish. Scientific Reports, 6 , 33726. 10.1038/srep33726 27641717
Davis, N. M. , Proctor, D. M. , Holmes, S. P. , Relman, D. A. , & Callahan, B. J. (2018). Simple statistical identification and removal of contaminant sequences in marker‐gene and metagenomics data. Microbiome, 6 (1 ), 226. 10.1186/s40168-018-0605-2 30558668
Earley, A. M. , Graves, C. L. , & Shiau, C. E. (2018). Critical role for a subset of intestinal macrophages in shaping gut microbiota in adult zebrafish. Cell Reports, 25 (2 ), 424–436. 10.1016/j.celrep.2018.09.025 30304682
El‐Gebali, S. , Mistry, J. , Bateman, A. , Eddy, S. R. , Luciani, A. , Potter, S. C. , Qureshi, M. , Richardson, L. J. , Salazar, G. A. , Smart, A. , Sonnhammer, E. L. L. , Hirsh, L. , Paladin, L. , Piovesan, D. , Tosatto, S. C. E. , & Finn, R. D. (2019). The Pfam protein families database in 2019. Nucleic Acids Research, 47 (D1 ), D427–D432. 10.1093/nar/gky995 30357350
Eren, A. M. , Esen, Ö. C. , Quince, C. , Vineis, J. H. , Morrison, H. G. , Sogin, M. L. , & Delmont, T. O. (2015). Anvi'o: An advanced analysis and visualization platform for ‘omics data. PeerJ, 3 , e1319. 10.7717/peerj.1319 26500826
Eren, A. M. , Kiefl, E. , Shaiber, A. , Veseli, I. , Miller, S. E. , Schechter, M. S. , Fink, I. , Pan, J. N. , Yousef, M. , Fogarty, E. C. , Trigodet, F. , Watson, A. R. , Esen, Ö. C. , Moore, R. M. , Clayssen, Q. , Lee, M. D. , Kivenson, V. , Graham, E. D. , Merrill, B. D. , … Willis, A. D. (2021). Community‐led, integrated, reproducible multi‐omics with anvi'o. Nature Microbiology, 6 (1 ), 3–6. 10.1038/s41564-020-00834-3
Ewels, P. , Ns Magnusson, M. , Lundin, S. , & Aller, M. K. (2016). MultiQC: Summarize analysis results for multiple tools and samples in a single report. Bioinformatics, 32 (19 ), 3047–3048. 10.1093/bioinformatics/btw354 27312411
Finn, R. D. , Clements, J. , & Eddy, S. R. (2011). HMMER web server: Interactive sequence similarity searching. Nucleic Acids Research, 39 , W29–W37. 10.1093/nar/gkr367 21593126
François‐Étienne, S. , Nicolas, L. , Eric, N. , Jaqueline, C. , Pierre‐Luc, M. , Sidki, B. , Aleicia, H. , Danilo, B. , Luis, V. A. , & Nicolas, D. (2023). Important role of endogenous microbial symbionts of fish gills in the challenging but highly biodiverse Amazonian blackwaters. Nature Communications, 14 (1 ), 3903. 10.1038/s41467-023-39461-x
Frøslev, T. G. , Kjøller, R. , Bruun, H. H. , Ejrnæs, R. , Brunbjerg, A. K. , Pietroni, C. , & Hansen, A. J. (2017). Algorithm for post‐clustering curation of DNA amplicon data yields reliable biodiversity estimates. Nature Communications, 8 (1 ), 1188. 10.1038/s41467-017-01312-x
Gaulke, C. A. , Barton, C. L. , Proffitt, S. , Tanguay, R. L. , & Sharpton, T. J. (2016). Triclosan exposure is associated with rapid restructuring of the microbiome in adult zebrafish. PLoS One, 11 (5 ), e0154632. 10.1371/journal.pone.0154632 27191725
Gaulke, C. A. , Beaver, L. M. , Armour, C. R. , Humphreys, I. R. , Barton, C. L. , Tanguay, R. L. , Ho, E. , & Sharpton, T. J. (2020). An integrated gene catalog of the zebrafish gut microbiome reveals significant homology with mammalian microbiomes . bioRxiv (p. 2020.06.15.153924) 10.1101/2020.06.15.153924
Gaulke, C. A. , Martins, M. L. , Watral, V. G. , Humphreys, I. R. , Spagnoli, S. T. , Kent, M. L. , & Sharpton, T. J. (2019). A longitudinal assessment of host‐microbe‐parasite interactions resolves the zebrafish gut microbiome's link to Pseudocapillaria tomentosa infection and pathology. Microbiome, 7 (1 ), 1–16. 10.1186/s40168-019-0622-9 30606251
Grunwald, D. J. , & Eisen, J. S. (2002). Headwaters of the zebrafish – emergence of a new model vertebrate. Nature Reviews Genetics, 3 , 717–724. 10.1038/nrg892
Gurbich, T. A. , Almeida, A. , Beracochea, M. , Burdett, T. , Burgin, J. , Cochrane, G. , Raj, S. , Richardson, L. , Rogers, A. B. , Sakharova, E. , Salazar, G. A. , & Finn, R. D. (2023). MGnify genomes: A resource for biome‐specific microbial genome catalogues. Journal of Molecular Biology, 435 (14 ), 168016. 10.1016/j.jmb.2023.168016 36806692
Hennersdorf, P. , Mrotzek, G. , Abdul‐Aziz, M. A. , & Saluz, H. P. (2016). Metagenomic analysis between free‐living and cultured Epinephelus fuscoguttatus under different environmental conditions in Indonesian waters. Marine Pollution Bulletin, 110 (2 ), 726–734. 10.1016/j.marpolbul.2016.05.009 27210559
Horz, H.‐P. , Scheer, S. , Huenger, F. , Vianna, M. E. , & Conrads, G. (2008). Selective isolation of bacterial DNA from human clinical specimens. Journal of Microbiological Methods, 72 (1 ), 98–102. 10.1016/j.mimet.2007.10.007 18053601
Hwang, W. Y. , Fu, Y. , Reyon, D. , Maeder, M. L. , Tsai, S. Q. , Sander, J. D. , Peterson, R. T. , Yeh, J. R. J. , & Joung, J. K. (2013). Efficient genome editing in zebrafish using a CRISPR‐Cas system. Nature Biotechnology, 31 (3 ), 227–229. 10.1038/nbt.2501
Hyatt, D. , Chen, G.‐L. , Locascio, P. F. , Land, M. L. , Larimer, F. W. , & Hauser, L. J. (2010). Prodigal: Prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics, 11 , 119. 10.1186/1471-2105-11-119 20211023
Kanehisa, M. , & Goto, S. (2000). KEGG: Kyoto encyclopedia of genes and genomes. Nucleic Acids Research, 28 (1 ), 27–30. 10.1093/nar/28.1.27 10592173
Kayani, M. U. R. , Yu, K. , Qiu, Y. , Shen, Y. , Gao, C. , Feng, R. , Zeng, X. , Wang, W. , Chen, L. , & Su, H. L. (2021). Environmental concentrations of antibiotics alter the zebrafish gut microbiome structure and potential functions. Environmental Pollution, 278 , 116760. 10.1016/j.envpol.2021.116760 33725532
Kayani, M. U. R. , Zaidi, S. S. A. , Feng, R. , Yu, K. , Qiu, Y. , Yu, X. , Chen, L. , & Huang, L. (2022). Genome‐resolved characterization of structure and potential functions of the zebrafish stool microbiome. Frontiers in Cellular and Infection Microbiology, 12 , 910766. 10.3389/fcimb.2022.910766 35782152
Le Doujet, T. , De Santi, C. , Klemetsen, T. , Hjerde, E. , Willassen, N.‐P. , & Haugen, P. (2019). Closely‐related Photobacterium strains comprise the majority of bacteria in the gut of migrating Atlantic cod (Gadus morhua). Microbiome, 7 (1 ), 64. 10.1186/s40168-019-0681-y 30995938
Lee, M. D. (2019). GToTree: A user‐friendly workflow for phylogenomics. Bioinformatics, 35 (20 ), 4162–4164. 10.1093/bioinformatics/btz188 30865266
Li, D. , Liu, C.‐M. , Luo, R. , Sadakane, K. , & Lam, T.‐W. (2015). MEGAHIT: An ultra‐fast single‐node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics, 31 (10 ), 1674–1676. 10.1093/bioinformatics/btv033 25609793
Li, H. (2018). Minimap2: Pairwise alignment for nucleotide sequences. Bioinformatics, 34 (18 ), 3094–3100. 10.1093/bioinformatics/bty191 29750242
Li, H. (2021). New strategies to improve minimap2 alignment accuracy. Bioinformatics, 37 (23 ), 4572–4574. 10.1093/bioinformatics/btab705 34623391
Limborg, M. T. , Chua, P. Y. S. , & Rasmussen, J. A. (2023). Unexpected fishy microbiomes. Nature Reviews. Microbiology, 21 (6 ), 346. 10.1038/s41579-023-00879-1
Marcos, S. , Parejo, M. , Estonba, A. , & Alberdi, A. (2022). Recovering high‐quality host genomes from gut metagenomic data through genotype imputation. Advances in Genetics, 3 (3 ), 2100065. 10.1002/ggn2.202100065
Marotz, C. A. , Sanders, J. G. , Zuniga, C. , Zaramela, L. S. , Knight, R. , & Zengler, K. (2018). Improving saliva shotgun metagenomics by chemical host DNA depletion. Microbiome, 6 (1 ), 42. 10.1186/s40168-018-0426-3 29482639
Marquet, M. , Zöllkau, J. , Pastuschek, J. , Viehweger, A. , Schleußner, E. , Makarewicz, O. , Pletz, M. W. , Ehricht, R. , & Brandt, C. (2022). Evaluation of microbiome enrichment and host DNA depletion in human vaginal samples using Oxford Nanopore's adaptive sequencing. Scientific Reports, 12 (1 ), 4000. 10.1038/s41598-022-08003-8 35256725
Martin, M. (2011). Cutadapt removes adapter sequences from high‐throughput sequencing reads. EMBnet.journal, 17 (1 ), 10–12. 10.14806/ej.17.1.200
McMurdie, P. J. , & Holmes, S. (2013). Phyloseq: An R package for reproducible interactive analysis and graphics of microbiome census data. PLoS One, 8 (4 ), e61217. 10.1371/journal.pone.0061217 23630581
Menzel, P. , Ng, K. L. , & Krogh, A. (2016). Fast and sensitive taxonomic classification for metagenomics with kaiju. Nature Communications, 7 , 11257. 10.1038/ncomms11257
Mikheenko, A. , Prjibelski, A. , Saveliev, V. , Antipov, D. , & Gurevich, A. (2018). Versatile genome assembly evaluation with QUAST‐LG. Bioinformatics, 34 (13 ), i142–i150. 10.1093/bioinformatics/bty266 29949969
Minich, J. J. , Härer, A. , Vechinski, J. , Frable, B. W. , Skelton, Z. R. , Kunselman, E. , Shane, M. A. , Perry, D. S. , Gonzalez, A. , McDonald, D. , Knight, R. , Michael, T. P. , & Allen, E. E. (2022). Host biology, ecology and the environment influence microbial biomass and diversity in 101 marine fish species. Nature Communications, 13 (1 ), 6978. 10.1038/s41467-022-34557-2
Oksanen, J. , Blanchet, F. G. , Friendly, M. , Kindt, R. , Legendre, P. , Mcglinn, D. , Minchin, P. R. , O'hara, R. B. , Simpson, G. L. , Solymos, P. , Henry, M. , Stevens, H. , Szoecs, E. , & Maintainer, H. W. (2020). Package “vegan” title community ecology package . Version 2.5‐7. https://CRAN.R‐project.org/package=vegan
Olm, M. R. , Brown, C. T. , Brooks, B. , & Banfield, J. F. (2017). dRep: A tool for fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de‐replication. The ISME Journal, 11 (12 ), 2864–2868. 10.1038/ismej.2017.126 28742071
Parks, D. H. , Chuvochina, M. , Waite, D. W. , Rinke, C. , Skarshewski, A. , Chaumeil, P.‐A. , & Hugenholtz, P. (2018). A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nature Biotechnology, 36 (10 ), 996–1004. 10.1038/nbt.4229
Parks, D. H. , Imelfort, M. , Skennerton, C. T. , Hugenholtz, P. , & Tyson, G. W. (2015). CheckM: Assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Research, 25 (7 ), 1043–1055. 10.1101/gr.186072.114 25977477
Pereira‐Marques, J. , Hout, A. , Ferreira, R. M. , Weber, M. , Pinto‐Ribeiro, I. , van Doorn, L.‐J. , Knetsch, C. W. , & Figueiredo, C. (2019). Impact of host DNA and sequencing depth on the taxonomic resolution of whole metagenome sequencing for microbiome analysis. Frontiers in Microbiology, 10 , 1277. 10.3389/fmicb.2019.01277 31244801
PubMed Search Results . (2023). PubMed. https://pubmed.ncbi.nlm.nih.gov/?term=zebrafish&filter=pubt.review&sort=date
Quast, C. , Pruesse, E. , Yilmaz, P. , Gerken, J. , Schweer, T. , Yarza, P. , Peplies, J. , & Glöckner, F. O. (2013). The SILVA ribosomal RNA gene database project: Improved data processing and web‐based tools. Nucleic Acids Research, 41 (D1 ), D590–D596. 10.1093/nar/gks1219 23193283
Rasmussen, J. A. , Kiilerich, P. , Madhun, A. S. , Waagbø, R. , Lock, E.‐J. R. , Madsen, L. , Gilbert, M. T. P. , Kristiansen, K. , & Limborg, M. T. (2023). Co‐diversification of an intestinal mycoplasma and its salmonid host. The ISME Journal, 17 (5 ), 682–692. 10.1038/s41396-023-01379-z 36807409
Rasmussen, J. A. , Villumsen, K. R. , Duchêne, D. A. , Puetz, L. C. , Delmont, T. O. , Sveier, H. , Jørgensen, L. V. G. , Præbel, K. , Martin, M. D. , Bojesen, A. M. , Gilbert, M. T. P. , Kristiansen, K. , & Limborg, M. T. (2021). Genome‐resolved metagenomics suggests a mutualistic relationship between mycoplasma and salmonid hosts. Communications Biology, 4 (1 ), 579. 10.1038/s42003-021-02105-1 33990699
Rawls, J. F. , Samuel, B. S. , & Gordon, J. I. (2004). Gnotobiotic zebrafish reveal evolutionarily conserved responses to the gut microbiota. Proceedings of the National Academy of Sciences of the United States of America, 101 (13 ), 4596–4601. 10.1073/pnas.0400706101 15070763
Richardson, L. , Allen, B. , Baldi, G. , Beracochea, M. , Bileschi, M. L. , Burdett, T. , Burgin, J. , Caballero‐Pérez, J. , Cochrane, G. , Colwell, L. J. , Curtis, T. , Escobar‐Zepeda, A. , Gurbich, T. A. , Kale, V. , Korobeynikov, A. , Raj, S. , Rogers, A. B. , Sakharova, E. , Sanchez, S. , … Finn, R. D. (2023). MGnify: The microbiome sequence data analysis resource in 2023. Nucleic Acids Research, 51 (D1 ), D753–D759. 10.1093/nar/gkac1080 36477304
Riiser, E. S. , Haverkamp, T. H. A. , Varadharajan, S. , Borgan, Ø. , Jakobsen, K. S. , Jentoft, S. , & Star, B. (2019). Switching on the light: Using metagenomic shotgun sequencing to characterize the intestinal microbiome of Atlantic cod. Environmental Microbiology, 21 (7 ), 2576–2594. 10.1111/1462-2920.14652 31091345
Roeselers, G. , Mittge, E. K. , Stephens, W. Z. , Parichy, D. M. , Cavanaugh, C. M. , Guillemin, K. , & Rawls, J. F. (2011). Evidence for a core gut microbiota in the zebrafish. The ISME Journal, 5 (10 ), 1595–1608. 10.1038/ismej.2011.38 21472014
Schneider, V. A. , Graves‐Lindsay, T. , Howe, K. , Bouk, N. , Chen, H.‐C. , Kitts, P. A. , Murphy, T. D. , Pruitt, K. D. , Thibaud‐Nissen, F. , Albracht, D. , Fulton, R. S. , Kremitzki, M. , Magrini, V. , Markovic, C. , McGrath, S. , Steinberg, K. M. , Auger, K. , Chow, W. , Collins, J. , … Church, D. M. (2017). Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Research, 27 (5 ), 849–864. 10.1101/gr.213611.116 28396521
Schubert, M. , Lindgreen, S. , & Orlando, L. (2016). AdapterRemoval v2: Rapid adapter trimming, identification, and read merging. BMC Research Notes, 9 , 88. 10.1186/s13104-016-1900-2 26868221
Sharpton, T. J. , Stagaman, K. , Sieler, M. J. , Arnold, H. K. , & Davis, E. W. (2021). Phylogenetic integration reveals the zebrafish core microbiome and its sensitivity to environmental exposures. Toxics, 9 (1 ), 1–15. 10.3390/toxics9010010
Shen, W. , Le, S. , Li, Y. , & Hu, F. (2016). SeqKit: A cross‐platform and ultrafast toolkit for FASTA/Q file manipulation. PLoS One, 11 (10 ), e0163962. 10.1371/journal.pone.0163962 27706213
Sieler, M. J., Jr. , Al‐Samarrie, C. E. , Kasschau, K. D. , Varga, Z. M. , Kent, M. L. , & Sharpton, T. J. (2023). Disentangling the link between zebrafish diet, gut microbiome succession, and Mycobacterium chelonae infection. Animal Microbiome, 5 (1 ), 38. 10.1186/s42523-023-00254-8 37563644
Spence, R. , Gerlach, G. , Lawrence, C. , & Smith, C. (2008). The behaviour and ecology of the zebrafish, Danio rerio . Biological Reviews of the Cambridge Philosophical Society, 83 (1 ), 13–34. 10.1111/j.1469-185X.2007.00030.x 18093234
Stagaman, K. , Sharpton, T. J. , & Guillemin, K. (2020). Zebrafish microbiome studies make waves. Lab Animal, 49 (7 ), 201–207. 10.1038/s41684-020-0573-6 32541907
Tatusov, R. L. , Fedorova, N. D. , Jackson, J. D. , Jacobs, A. R. , Kiryutin, B. , Koonin, E. V. , Krylov, D. M. , Mazumder, R. , Mekhedov, S. L. , Nikolskaya, A. N. , Rao, B. S. , Smirnov, S. , Sverdlov, A. V. , Vasudevan, S. , Wolf, Y. I. , Yin, J. J. , & Natale, D. A. (2003). The COG database: An updated version includes eukaryotes. BMC Bioinformatics, 4 , 41. 10.1186/1471-2105-4-41 12969510
Tennant, M. D. , Zerucha, T. , & Bouldin, C. M. (2019). Cost‐efficient and effective mutagenesis in Zebrafish with CRISPR/Cas9. Eastern Biologist Special Issue, 1 , 64–74.
Thormar, E. A. (2023a). MAGs_summary. Figshare . Dataset. 10.6084/m9.figshare.24866772.v1
Thormar, E. A. (2023b). ALL_MAGs. Figshare . Dataset. 10.6084/m9.figshare.24866790.v1
Thormar, E. A. (2023c). DREP. Figshare . Dataset. 10.6084/m9.figshare.24866841.v1
Thormar, E. A. , Hansen, S. B. , von Jørgensen, L. G. , & Limborg, M. T. (2023). PRJEB71469.European nucleotide archive (ENA) . [dataset]. https://www.ebi.ac.uk/ena/browser/view/PRJEB71469
Wickham, H. (2016). ggplot2: Elegant graphics for data analysis. Springer. https://play.google.com/store/books/details?id=XgFkDAAAQBAJ
Xu, L. , Xu, Y. , Xue, T. , Zhang, X. , & Li, J. (2021). AdImpute: An imputation method for single‐cell RNA‐Seq data based on semi‐supervised autoencoders. Frontiers in Genetics, 12 , 739677. 10.3389/fgene.2021.739677 34567089
Yap, M. , Feehily, C. , Walsh, C. J. , Fenelon, M. , Murphy, E. F. , McAuliffe, F. M. , van Sinderen, D. , O'Toole, P. W. , O'Sullivan, O. , & Cotter, P. D. (2020). Evaluation of methods for the reduction of contaminating host reads when performing shotgun metagenomic sequencing of the milk microbiome. Scientific Reports, 10 (1 ), 21665. 10.1038/s41598-020-78773-6 33303873
Yu, Y. , Lee, C. , Kim, J. , & Hwang, S. (2005). Group‐specific primer and probe sets to detect methanogenic communities using quantitative real‐time polymerase chain reaction. Biotechnology and Bioengineering, 89 (6 ), 670–679. 10.1002/bit.20347 15696537
Zhang, Z. , Peng, Q. , Huo, D. , Jiang, S. , Ma, C. , Chang, H. , Chen, K. , Li, C. , Pan, Y. , & Zhang, J. (2021). Melatonin regulates the neurotransmitter secretion disorder induced by caffeine through the microbiota‐gut‐brain Axis in zebrafish (Danio rerio). Frontiers in Cell and Developmental Biology, 9 , 678190. 10.3389/fcell.2021.678190 34095150
