
==== Front
J Biol Chem
J Biol Chem
The Journal of Biological Chemistry
0021-9258
1083-351X
American Society for Biochemistry and Molecular Biology

S0021-9258(24)02126-4
10.1016/j.jbc.2024.107625
107625
Research Article
The molecular basis of cereal mixed-linkage β-glucan utilization by the human gut bacterium Segatella copri
Golisch Benedikt 12‡
Cordeiro Rosa Lorizolla 1‡
Fraser Alexander S.C. 1
Briggs Jonathon 1
Stewart William A. 13
Van Petegem Filip 3
Brumer Harry brumer@msl.ubc.ca
1234∗
1 Michael Smith Laboratories, University of British Columbia, Vancouver, British Columbia, Canada
2 Department of Chemistry, University of British Columbia, Vancouver, British Columbia, Canada
3 Department of Biochemistry and Molecular Biology, University of British Columbia, Vancouver, British Columbia, Canada
4 Department of Botany, University of British Columbia, Vancouver, British Columbia, Canada
∗ For correspondence: Harry Brumer brumer@msl.ubc.ca
‡ These authors contributed equally to this work.

08 8 2024
9 2024
08 8 2024
300 9 10762517 5 2024
15 7 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Mixed-linkage β(1,3)/β(1,4)-glucan (MLG) is abundant in the human diet through the ingestion of cereal grains and is widely associated with healthful effects on metabolism and cholesterol levels. MLG is also a major source of fermentable glucose for the human gut microbiota (HGM). Bacteria from the family Prevotellaceae are highly represented in the HGM of individuals who eat plant-rich diets, including certain indigenous people and vegetarians in postindustrial societies. Here, we have defined and functionally characterized an exemplar Prevotellaceae MLG polysaccharide utilization locus (MLG-PUL) in the type-strain Segatella copri (syn. Prevotella copri) DSM 18205 through transcriptomic, biochemical, and structural biological approaches. In particular, structure-function analysis of the cell-surface glycan-binding proteins and glycoside hydrolases of the S. copri MLG-PUL revealed the molecular basis for glycan capture and saccharification. Notably, syntenic MLG-PULs from human gut, human oral, and ruminant gut Prevotellaceae are distinguished from their counterparts in Bacteroidaceae by the presence of a β(1,3)-specific endo-glucanase from glycoside hydrolase family 5, subfamily 4 (GH5_4) that initiates MLG backbone cleavage. The definition of a family of homologous MLG-PULs in individual species enabled a survey of nearly 2000 human fecal microbiomes using these genes as molecular markers, which revealed global population-specific distributions of Bacteroidaceae- and Prevotellaceae-mediated MLG utilization. Altogether, the data presented here provide new insight into the molecular basis of β-glucan metabolism in the HGM, as a basis for informing the development of approaches to improve the nutrition and health of humans and other animals.

Keywords

human gut microbiota
microbiome
glycoside hydrolase
glycan-binding protein
TonB-dependent transporter
Abbreviations

AGE affinity gel electrophoresis

BCA bicinchoninic acid

BSA bovine serum albumin

CAZyme carbohydrate-active enzyme

GM glucomannan

HEC hydroxyethyl cellulose

Hepes 4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid

HGM human gut microbiota

HPAEC-PAD high-performance anionic-exchange chromatography with pulsed amperometric detection

HTCS hybrid two-component system

IMAC immobilized metal ion affinity chromatography

ITC isothermal titration calorimetry

MLG mixed-linkage β(1,3)/β(1,4)-glucan

pNP para-nitrophenol

PUL polysaccharide utilization locus

SGBP cell-surface glycan-binding protein

TBDT TonB-dependent transporter

XyG xyloglucan

Reviewed by members of the JBC Editorial Board. Edited by Chris Whitfield
==== Body
pmcThe human gastrointestinal tract is home to trillions of microorganisms collectively known as the human gut microbiota (HGM) (1). The homeostasis and metabolism of the HGM is central to human nutrition and health (2, 3, 4). For example, we are critically dependent on the HGM for the catabolism of the complex carbohydrates that constitute dietary fiber, due to a paucity of digestive glycosidases encoded by our own genomes (5). In contrast, many bacteria of the HGM possess diverse arsenals of carbohydrate-active enzymes (CAZymes) that enable the saccharification of a wide range of dietary glycans (6). Subsequent fermentation of the released monosaccharides produces short-chain fatty acids and other beneficial metabolites. Short-chain fatty acids, in particular, serve as the primary energy source for colonocytes, contribute to up to 10% of daily caloric intake, and modulate postprandial glucose levels, food intake, and energy expenditure (4, 7, 8). Also crucial, a diverse and well-balanced HGM ecosystem forms a natural ecological barrier to the invasion of the gut by pathogenic bacteria, fungi, and viruses (4, 9).

Within individuals, the composition of the HGM is variable and dynamic, influenced by numerous host-related and environmental factors, including host genetics, sex, age, stress, and especially diet (4, 10). The HGM is, in general, dominated by approximately equal proportions of bacteria from the phyla Bacteroidota (synonym Bacteroidetes) and Bacillota (syn. Firmicutes), with Actinobacteria, Proteobacteria, Fusobacteria, and Verrucomicrobia together constituting up to 10% of other bacteria (4). Numerous studies have recognized the importance of high taxa diversity, high microbial gene richness, and stable “functional cores” as beneficial for human health (10, 11). Yet due to its inherent complexity and individual variation, definition of a “healthy HGM” remains elusive (4, 12).

Large differences in HGM composition have been observed, particularly between individuals who consume postindustrial (“Western”) diets dominated by starch, protein, fat, and processed foods and those who consume diets rich in plant cell walls, including traditional hunter-gatherer and agrarian populations, as well as Western vegetarians and vegans (10, 13, 14). These groups differ notably in genera from the phylum Bacteroidota, which represents ca. 50% of HGM taxa: Bacteroides species (family Bacteroidaceae) predominate in postindustrial HGM, whereas Prevotella species (recently reclassified as Segatella and related genera in the family Prevotellaceae (15, 16)) dominate in predominant vegetarians. This distinction results in segregation into distinct “enterotypes” (11, 17). Strikingly, replacement of Prevotellaceae with Bacteroidaceae in individual HGM is associated with migration from Asia to North America (18). Prevotellaceae are classically known as key members of animal rumens (19), and it is only through this recent shift in microbiome analyses toward a more global outlook that their ubiquity in human (monogastric) HGMs has been revealed (16, 17, 20).

Due to this initial focus on North American and European HGM, we currently have an extensive understanding of the molecular systems that enable complex glycan utilization in human gut Bacteroidaceae. Bacteroides species encode entire suites of proteins and enzymes required for the sensing, capture, transport, and deconstruction of individual complex carbohydrates in contiguous, multi-gene polysaccharide utilization loci (PULs) (21). Bacteroides PULs specific for individual dietary plant, algal, animal, and microbial glycans have been characterized through combined microbiological, biochemical, and structural biological studies (reviewed in (22, 23, 24), see (25, 26, 27, 28, 29, 30) for recent examples). Prevotellaceae from the HGM are likewise predicted to encode extensive cohorts of CAZymes in PULs (5, 31, 32, 33). However, our experimental knowledge of PUL structure and function in this taxonomic family is currently limited to the partial characterization of two plant β-xylan PULs in the human fecal isolate Segatella copri (syn. Prevotella copri) (34).

In our present study, we have extended our knowledge of complex carbohydrate metabolism in Segatella through the identification and characterization of a PUL responsible for the utilization of mixed-linkage β(1,3)/β(1,4)-glucan (MLG, Fig. 1) in the type-strain S. copri CB7 (syn. DSM18205; hereafter, S. copri) (35). MLG is ubiquitous and abundant in the human diet via cereal grains, including wheat, barley, rice, maize, oats, and rye (36) and thus constitutes a major source of prebiotic glucose for the HGM. Moreover, dietary MLG has demonstrable effects on human health, including lowering cholesterol and blood glucose levels, reducing insulin resistance, and mitigating metabolic syndrome (36, 37, 38). Our present data provide a validated reference for the bioinformatic prediction of MLG utilization in Segatella genomes and gut metagenomes.Figure 1 Identification of a mixed-linkage β-glucan polysaccharide utilization locus (MLG-PUL) in Segatella copri.A, general structure of cereal MLG, comprised primarily of cellotriosyl and cellotetraosyl units connected by β(1,3) linkages. B, growth of S. copri on glucose and MLG. Error bars represent SEM (n = 3) at each sampling timepoint. C, volcano plot identifying 324 differentially regulated genes when grown on MLG versus glucose. Other S. copri genes did not clear inclusion thresholds and were omitted. Purple data points are part of the putative MLG-PUL.

Results and discussion

Identification of a putative MLG-PUL in S. copri

To identify potential MLG utilization genes in S. copri, we performed comparative total RNA-seq. Concordant with previous results (32), S. copri grew in minimal medium containing either glucose or barley MLG as sole carbon sources, albeit with a longer lag phase for the latter (Fig. 1B). Differential transcript analysis revealed the upregulation of a set of colocalized genes in response to MLG (Fig. 1C), constituting an apparent PUL (Fig. 2A).Figure 2 Segatella copri MLG-PUL, homology, and biochemical model.A, putative S. copri MLG-PUL and syntenic PULs in selected Prevotellaceae and Bacteroidaceae species. Species names and synonyms are according to the List of Prokaryotic Names with Standing in Nomenclature (97) and recent reclassification of Prevotella species (15, 16). Genome accessions, gene locus tags, and annotations are provided for each gene as Supporting Information in a Tab Separated Value file. B, working model of the MLG utilization system in the cell envelope. Protein localization is predicted based on signal peptide analysis with SignalP 6.0 (48). GHn, Glycoside Hydrolase family members; HTCS, Hybrid Two-Component System sensor/regulator; MFS, Major Facilitator Superfamily transporter; SGBP, cell-Surface Glycan-Binding Proteins; SusR, Starch Utilization System gene R; TBDT, TonB-Dependent Transporter.

Protein sequence analysis indicated that these genes include the tandem susC/susD homolog pair, encoding a TonB-dependent transporter (TBDT, locus tag NQ544_01015) and an associated cell-surface glycan-binding protein (SGBP)-A lid protein (NQ544_01020), respectively, which is a hallmark of canonical Bacteroides PULs (31, 39, 40). The putative MLG-PUL also contained genes encoding members of glycoside hydrolase families GH5-subfamily 4 (GH5_4, predicted endo-(xylo)glucanase; NQ544_01010) and GH3 (predicted exo-β-glucosidase, NQ544_01030) (41). By analogy with other PULs (21, 24, 42), a gene encoding a predicted, nonhomologous SGBP-B (SGBP-B, NQ544_01025) was located immediately downstream from the susD homolog. Lastly, a nonupregulated gene encoding a predicted hybrid two-component system (HTCS, NQ544_01005) sensor/regulator (39) was found upstream of the GH5_4-encoding gene (Fig. 2A). This ensemble of genes was found to be conserved in other bacteria from the family Prevotellaceae, including Segatella, Prevotella, Hoylesella, and Xylanibacter species. However, despite overall synteny, there was poor sequence similarity with proteins in Bacteroides MLG-PULs (exemplified by Bacteroides ovatus). Most notably, this included the apparent replacement of the GH16 mixed-linkage endo-β(1,3)/β(1,4)-glucanase in Bacteroides species (43) with the GH5_4 member in S. copri (Fig. 2A). This observation introduced ambiguity regarding the biological function of this PUL because members of GH5_4, including those from human gut Bacteroides species, are commonly associated with endo-xyloglucanase activity (44, 45, 46, 47).

We also observed that NQ544_01055, encoding a GH94 member (predicted cellobiose phosphorylase), was slightly but significantly upregulated during growth on MLG versus glucose (Fig. 1). The gene was not formally considered to be part of the MLG-PUL because it is located ca. 6 kilobases away from the final GH3-encoding gene and on the opposite DNA strand of the predicted PUL (Fig. 2A). However, its coregulation suggested a potential accessory role in MLG utilization (vide infra).

Signal peptide analysis using SignalP 6.0 (48) predicted that the GH5_4, SGBP-A, and SGBP-B proteins were all likely to be cell-surfaced localized following N-terminal cysteine lipidation and signal peptidase II cleavage (49). The GH3 member possessed a signal peptidase I cleavage site, suggesting periplasmic localization, whereas the GH94 member lacked a predicted signal peptide, indicating cytoplasmic localization. Together, this analysis suggested a working model of the putative MLG-PUL and GH94 member, by analogy with other PUL systems (21, 42) (Fig. 2B).

SGBP-A and SGBP-B recognize MLG through distinct protein architectures

SGBP-A is a SusD homolog (42), which is predicted to interact directly with the TBDT (a SusC homolog (50)) to form a flexible lid over the β-barrel to gate entry of polysaccharide breakdown products ((51) and references therein). SGBP-B, on the other hand, is a member of a group of sequence-diverse, glycan-binding proteins that are primarily unified by their common genomic position downstream of the SusD homolog in PULs (sometimes referred to as “SusE-positioned”) (24, 42). Both proteins were successfully produced recombinantly as constructs lacking their predicted native signal peptides and disordered N-terminal peptide linkers (Figs. S1 and S2). To validate the role of these proteins in MLG capture at the cell surface, we first tested the binding of a library of soluble polysaccharides using affinity gel electrophoresis (AGE, Fig. 3). Both proteins bound strongly to MLG but notably also to other polysaccharides containing β(1,4)-glucosyl backbone linkages, viz. xyloglucan (XyG), glucomannan (GM), and hydroxyethyl cellulose (HEC). On the other hand, neither SGBP bound to β(1,3)-glucans, β-mannans, β(1,4)-galactan, nor β(1,4)-xylan.Figure 3 Binding of SGBPs to polysaccharides analyzed by affinity gel electrophoresis. Gels were loaded with SGBP-AP50-K595 (SGBP-A) and SGBP-BF42-E332 (SGBP-B); in select cases, SGBP-AP50-K595 after TEV-cleavage of the hexahistidine purification tag (SGBP-A_TEV) is included. Bovine serum albumin (BSA) was included as a noninteracting standard and was observed both in monomeric and dimeric forms. Gels were run on separate occasions and are grouped accordingly, including individual control gels without added polysaccharide (water-only). Polysaccharide concentrations in the gel are given as w/v.

Isothermal titration calorimetry (ITC) validated the polysaccharide binding observed by AGE and indicated decreasing binding affinity in the order MLG > XyG > GM >> HEC (Fig. S3, Table S1). Notably, SGBP-A displayed a ca. 10-fold higher binding affinity for MLG over XyG and GM, whereas SGBP-B bound MLG and XyG with comparable affinity (slightly less for GM). In light of the working model (Fig. 2), these data suggest that binding of S. copri to plant digesta, mediated by SGBP-B, is promiscuous, whereas import through the TBDT, gated by SGBP-A (a SusD homolog), is more selective. Notably, the binding of both proteins to cellotetraose was undetectable by ITC, suggesting that longer oligosaccharides and polysaccharides serve as natural substrates for both of these SGBPs, mirroring results observed for SGBP-A and -B from the B. ovatus MLG-PUL (52).

At the cell surface, SGBP-B is likely to play a role in initiating binding of the bacterium to plant digesta and/or attracting soluble polysaccharides to the cell surface (42, 52, 53). Indeed, SGBP-B is predicted to comprise three domains joined by short linkers (Fig. S2), anchored by N-terminal lipidation. Unfortunately, we were unable to generate an experimental tertiary structure due to poor crystal diffraction, despite extensive screening. An Alphafold model of SGBP-B showed that three aromatic amino acid residues are exposed to the solvent of the C-terminal domain, forming a flat surface that could accommodate the backbone of β(1,4)-glucans (Fig. 4). Indeed, AGE of the individual modules of the SGBP-B showed that only the C-terminal module participates in polysaccharide binding (Fig. S2D). Superposition with the experimental structures of the SGBP-B from the B. ovatus MLG-PUL, in complex with MLG- and cello-oligosaccharides (PDB 6E57 and 6E9B, respectively (52)), confirmed the overall structural homology of the binding face, despite significant differences in primary structure. The S. copri MLG-PUL SGBP-B has a shorter and flatter binding face, due to the lack of a fourth aromatic residue and a loop truncation, which may explain its greater promiscuity toward MLG and XyG than the B. ovatus MLG-PUL homolog (Table S1 cf. (52)).Figure 4 Structural model of SGBP-B.A, Alphafold model of SGBP-B represented in surface (white) and cartoon, colored according to the confidence score (pLDDT) for each residue (very low pLDDT: <50, very high pLDDT: >90). The N terminus of the protein is predicted to contain an N-terminal cysteine lipidation site for membrane anchoring. B, C-terminal domain-binding surface, showing three tryptophan residues comprising an aromatic binding platform (magenta). A nonconserved aspartic acid is located near the binding site, which could additional hydrogen-bonding interaction with substrates. C, structure-guided protein sequence alignment of the C-terminal–binding domains of the Segatella copri and Bacteroides ovatus MLG-PUL SGBP-B homologs. Key binding residues are highlighted in magenta. D, superposition of the S. copri SGBP-B model with experimental complexes (52) of the B. ovatus homolog with cellopentaose (PDB 6E57) and MLG-hexaose (PDB 6E9B). Oligosaccharides are represented in sticks with green carbons.

To investigate the molecular basis of MLG specificity in the SusD homolog SGBP-A, we crystallized and solved the structures for the free protein and in complex with cellopentaose, at 2.00 Å and 1.50 Å resolution, respectively (Table S2). The structures were determined by molecular replacement using an Alphafold model as the search coordinates in the space group P212121 with Rwork/Rfree of 18.5%/22.5% for free SGBP-A and of 15.2%/16.6% for the cellopentaose complex (Fig. 5). Both structures displayed the characteristic SusD-like protein fold (54), comprising structurally conserved tetratricopeptide repeat motifs (PFAM PF07980) and a variable glycan-binding region (24, 55). Notably, no electron density was observed in the free form for a nonconserved loop inserted in the glycan-binding region that comprises ten amino acid residues (Thr267 to Asp276), demonstrating intrinsic disorder of this region. In the cellopentaose-bound structure, this loop is stabilized and clearly resolved due to direct interactions between Tyr269 and the oligosaccharide (Fig. 5).Figure 5 Crystallography of SGBP-A.A, SusD-like overall fold of SGBP-A in the free form. The tetratricopeptide repeat (TPR) domains, a pair of conserved α-helices (∗), variable regions associated with glycan binding in diverse homologs are indicated. B, complex with cellopentaose (B), including tetratricopeptide repeat (TPR) domains, a pair of conserved α-helices (∗), the variable regions that is associated with glycan binding in diverse homologs. C, schematic representation indicating the position of the TPR motifs and the flexible loop in the primary structure. The dashed box indicates the N-terminal region not observed in the structures, likely due to intrinsic flexibility. D, interactions at the binding interface with cellopentaose. E, superposition with Bacteroides ovatus SGBP-A:MLG-heptasaccharide complex (PDB 6E61, (52)), showing the residues composing the aromatic platform and MLG-heptasaccharide.

Specifically, the aromatic ring of Tyr269 stacks with the reducing-end glucosyl residue, while the phenolic hydroxyl is within hydrogen-bonding distance to O-2 of the second glucosyl residue. A second loop (Gln289 to Ser292) also undergoes a conformational change upon oligosaccharide binding, thereby allowing additional hydrogen bonding. The complex also revealed interactions with a conserved aromatic platform, comprising Trp86, Tyr363, and Trp336. Trp86 and Tyr363 were directly observed to stack with the middle and nonreducing end of cellopentaose, which was otherwise too short to span Trp366. However, superposition with the SGBP-A homolog in the B. ovatus MLG-PUL, in complex with MLG-heptaose (PDB 6E61, (52)), suggests that Trp366 likewise provides a stacking interaction at kinked positions in the MLG backbone. Substrate binding is further complemented by a series of direct and water-mediated hydrogen bonds (Fig. 5).

To gain additional insight into the likely role of the SusD homolog SGBP-A in carbohydrate transport, we modeled it in complex with an Alphafold model of the associated TBDT (SusC homolog), using a homologous complex from the Bacteroides thetaiotaomicron levan-PUL (PDB 8AA1, (51)) as a reference (Fig. S4). Interestingly, the model suggests that the S. copri TBDT possesses a groove containing aromatic residues that could bind to MLG oligosaccharides. In the closed conformation, the flexible loop from the MLG-PUL SGBP-A (vide supra) fits into the transporter pore, near two aromatic residues (Tyr477 and Trp832) of the TBDT groove. Thus, the flexibility in this region could be an adaptation to facilitate glycan transfer from the SGBP-A to the transporter. Notably, the cavity formed between the SGBP and TBDT could allow the accommodation of long oligosaccharides (DP> 7) in the closed complex, which agrees with our observations that SGBP-A prefers longer oligosaccharides.

GH5_4 is a specific MLGase that preferentially cleaves β-1,3-linkages

The predicted GH5_4 preprotein included a 20 amino acid lipoprotein signal peptide with a signal peptidase II cleavage site before Cys21 and a 20 amino acid–predicted disordered region before the GH5_4 module. Notably, the protein lacked the N-terminal spacer domain (PFAM PF13004) that is sometimes found in Bacteroides cell-surface GHs, including the GH5_4 endo-xyloglucanase from the B. ovatus XyG-PUL (46, 56). A protein construct for the predicted mature protein, GH5_4C21-E406, was recombinantly produced and purified for enzymology (Fig. S5).

An initial substrate screen showed that GH5_4 is a strict mixed-linkage endo-β(1,3)/β(1,4)-glucanase with high specific activity on barley MLG (Fig. S6). The specific activities on the next best substrates, XyG and GM, were ca. 40- and 80-fold lower, respectively, while specific activities on the artificial soluble β(1,4)-glucans carboxymethylcellulose and HEC were over 400-fold lower than on MLG. Activity on the highly branched bacterial β(1,4)-glucan xanthan, a common food additive (26), was undetectable, and activities on β(1,3)-glucans and β(1,4)-xylan were barely detectable. Further analysis on MLG suggested that GH5_4 was maximally active in the pH range 4.0 to 5.5 and up to 50 °C (Fig. S7).

Initial-rate kinetic analysis further confirmed that MLG is the preferred substrate of GH5_4, with a selectivity constant (kcat/Km value) 560-fold higher than for xyloglucan, due to ca. 70-fold higher kcat value and a ca. 8-fold lower Km value (Fig. 6, Table S3). Michaelis–Menten analysis also showed that GM was an equally poor substrate as XyG and that carboxymethylcellulose and HEC were worse still, concordant with the specific activity data (Fig. S6). These results are notable for two reasons: first, GH5_4 is commonly associated with endo-xyloglucanase activity (44), so the kinetics data resolves this ambiguity and suggests that the specificity of the PUL is mediated in large part by this endo-hydrolase at the cell surface. Second, poor activity on substrates containing only β(1,4)-glucosidic bonds suggested a limited ability to cleave these linkages.Figure 6 Initial-rate kinetics of the hydrolysis of polysaccharides containing backbone β(1,4)-glucosyl residues by GH5_4C21-E406.A, mixed-linkage β(1,3)/β(1,4)-glucan. B, xyloglucan (●), glucomannan (▲), hydroxyethyl cellulose (▼), and carboxymethyl cellulose (♦). Error bars represent SDs from the mean values (n = 3). Lines represent fits of the standard Michaelis–Menten equation to the data.

Indeed, product analysis by HPLC subsequently showed that GH5_4 preferentially hydrolyzed the β(1,3) linkages in MLG (Fig. 1) with an endo-dissociative mode of action (for a definition, see (57, 58)). GH5_4 initially generated a series of longer oligosaccharides which were subsequently converted over time (Fig. S8). Cellotriose (G4G4G) and cellotetraose (G4G4G4G) were the shortest, predominant oligosaccharides observed at early to middle time points (2 min–1 h). Notably, the corresponding mixed-linkage gluco-oligosaccharides G4G3G and G4G4G3G, which are typically produced through hydrolysis of β(1,4) linkages by canonical GH16 MLGases, were not observed. Upon extended incubation (2 h–45 h), cellotetraose was nonetheless hydrolyzed to cellobiose (G4G). In addition to cellobiose and cellotriose, the limit-digest also contained a lesser amount of the mixed-linkage tetrasaccharide G3G4G4G, all of which were not appreciably hydrolyzed by the GH5_4 MLGase (a trace amount of glucose and no laminaribiose, G3G, were observed at 45 h).

To understand further GH5_4 active-site selectivity, we incubated the enzyme with individual cellooligosaccharides [β(1,4)], laminarin oligosaccharides [β(1,3)], and MLG oligosaccharides [β(1,3)/β(1,4)] (Fig. S9, summarized in Fig. 7). Cellobiose (G4G) and cellotriose (G4G4G) were not hydrolyzed, as observed in the MLG limit-digestion (Fig. S8). Likewise, other di- and tri-saccharides were not hydrolyzed i.e., laminaribiose (G3G), laminaritriose (G3G3G), G3G4G, and G4G3G (trace G4G and Glc were observed with G4G3G). Also as observed in the MLG limit-digestion (Fig. S8), cellotetraose (G4G4G4G) was cleaved symmetrically to yield cellobiose (G4G), with very limited asymmetric cleavage to produce cellotriose (G4G4G) and glucose (G). Laminaritetraose (G3G3G3G) was not hydrolyzed by GH5_4, nor was G3G4G4G, a low-abundance MLG limit-digestion product (Fig. S9 cf. Fig. S8). However, G4G3G4G was readily hydrolyzed to cellobiose (2× G4G) by cleavage of the central β(1,3) linkage. Moreover, G4G4G3G was hydrolyzed to nearly equal proportions of cellotriose (G4G4G) and glucose through β(1,3) cleavage and cellobiose (G4G) and laminaribiose (G3G) through β(1,4) cleavage. Finally, cellopentaose (G4G4G4G4G) was hydrolyzed by GH5_4 into cellotriose (G4G4G) and cellobiose (G4G). Performing the enzyme-catalyzed hydrolysis of cellopentaose in 18O-enriched water indicated that cleavage occurred at both internal glucosidic bonds, with a ca. 4:1 preference for G4G4G + G4G cleavage over G4G + G4G4G cleavage (Fig. S10). Considered together, these data allowed us to derive an initial map of subsite specificity within the GH5_4 active site (Fig. 7). Overall, recognition of at least four glucosyl residues is required for efficient catalysis, and β(1,3) linkages are not tolerated between glucosyl residues in subsites −3 to −2 and −2 to −1 (subsite nomenclature according to (59)). The data also suggest the predominance of at least three negative and two positive subsites.Figure 7 Oligosaccharide specificity of GH5_4C21-E406.A, summary of data based on qualitative analysis based on product analysis by HPAEC-PAD (Fig. S8). Potential enzyme subsites are numbered according to (59). Catalytically productive binding modes are indicated with a green checkmark. B, initial-rate kinetic analysis of the hydrolysis of selected oligosaccharides. The time-dependent appearance of products characteristic of individual subsite-binding modes was monitored by HPAEC-PAD. The lower panel shows an expansion of the graph for the poorest substrates. The following hydrolysis reactions were monitored: G4G4G4G4G → G4G4G + G4G (▼), G4G3G4G → 2 G4G (►), G4G4G4G → 2 G4G (▪), G4G4G3G → G4G + G3G (♦), G4G4G4G → G4G4G + G (●), and G4G4G3G → G4G4G + G (◄). Error bars represent SDs from the mean values (n = 3). Lines represent fits of the standard Michaelis–Menten equation to the data.

To determine the contributions of individual subsites to catalysis, we monitored initial-rate kinetics of oligosaccharide hydrolysis by HPLC (60) (Fig. 7). Michaelis–Menten analysis indicated that the selectively of GH5_4 for the central β(1,3) linkage in the MLG tetrasaccharide G4G3G4G was ca. 6.5-fold greater than for the corresponding β(1,4) linkage in cellotetraose (G4G4G4G), based on kcat/Km values (Table S4). With comparable Km values, this difference indicates that kcat effects primarily dictate the observed specificity of the enzyme for β(1,3) linkages in the hydrolysis of the MLG polysaccharide (Fig. S8 cf. Fig. 1). G4G3G4G was hydrolyzed exclusively to two molecules of cellobiose through binding in subsites −2 to +2. On the other hand, cellotetraose (G4G4G4G) was simultaneously observed to be hydrolyzed in both a −2 to +2 binding mode (yielding two molecules of cellobiose, G4G) and an apparent 3 to +1 binding mode (yielding cellotriose, G4G4G, plus glucose, G), albeit with a 6.5-fold lower selectivity (Table S4). This reflects a significant free-energy contribution of the +2 subsite to binding (ΔΔG = – RT ln[(kcat/Km)1/(kcat/Km)2 = −4.9 kJ/mol), which is not compensated by binding of a glucosyl residue in subsite −3.

Notably, cellopentaose (G4G4G4G4G) is hydrolyzed to G4G4G + G4G with a kcat/Km value ca. 27-fold greater than cellotetraose in −2 to +2 mode and 180-fold greater than cellotetraose in −3 to +1 mode. These values would equate to free-energy contributions of −8.6 kJ/mol and −13 kJ/mol for the −3 and +2 subsites, respectively. However, considering that cellopentaose is hydrolyzed through two modes i.e, G4G4G|4G4G and G4G|4G4G4G, in a ca. 4:1 ratio (Fig. S10, vide supra), these values should be scaled accordingly. We also note that crystallography of the GH5_4 enzyme indicates that the active site does not contain a +3 subsite (vide infra). It is therefore notable that despite having a less-preferred β(1,4) scissile bond, the contributions of the additional subsite (−3) make cellopentaose (G4G4G4G4G) a 4-fold better substrate (based on kcat/Km values) than the MLG tetrasaccharide G4G3G4G. In addition to cellotriosyl and cellotetraosyl units separated by β(1,3) bonds, the MLG polysaccharide can also contain small amounts of cellopentaosyl units depending on the source (36). These kinetic data indicate that this longer cellooligosaccharide unit would also be readily hydrolyzed by the GH5_4 at the S. copri cell surface.

Considering the initial-rate kinetic data for the MLG tetrasaccharides G3G4G4G and G4G4G3G, some notable observations arise. We see immediately that a β(1,3) linkage is not tolerated between subsites −2 and −1 nor −3 and −2. Neither β(1,4) bond in G3G4G4G is hydrolyzed. Additionally, the impact of the loss of binding at subsite −2 is so great that hydrolysis of the otherwise preferred β(1,3) bond likewise does not occur in a −1 to +3 mode. In contrast, G4G4G3G is hydrolyzed in both the −2 to +2 mode (yielding G4G + G3G) and the −3 to +1 mode (yielding G4G4G + G) with approximately equivalent kcat/Km values, which are nonetheless over 20-fold lower than G4G3G4G. Thus, a β(1,3) linkage spanning subsites +1 to +2 is not vigorously excluded (and in fact, results in a slight −1.9 kJ/mol gain over G4G4G4G). However, positioning the β(1,3) linkage for hydrolysis, plus the additional binding contribution of subsite −3, nearly compensates for the loss of binding in subsite +2. Finally, the inactivity of the GH5_4 on all-β(1,3)-glucans (curdlan, laminarin, yeast β-glucan, Fig. 7) is explained by the lack of activity on laminaritetraose (G3G3G3G). The two binding modes that place β(1,3) linkages spanning negative subsites are excluded, as is the potential −1 to +3 mode.

To gain insights into the molecular basis of substrate selectivity of the GH5_4 member, we solved the structures of the free protein and the Glu331Ala mutant in complex with cellotriose at 1.6 Å (space group P21, Rfree/Rwork = 18.80%/21.35%) and 2.11 Å (space group P3221, Rfree/Rwork = 21.85%/24.42%) resolution, respectively (Table S6). Four molecules were found in the asymmetric unit of the free structure, while two were observed for the GH5_4:cellotriose complex, sharing an RMSD of 0.089 to 0.115 Å over 307 Cα atoms (Fig. S10). However, size-exclusion chromatography indicated that this protein is monomeric in solution (data not shown), concordant with PISA analysis (61). The tertiary structure comprises the conserved (α/β)8-barrel fold of Clan GH-A and two glutamates (Glu197 and Glu331, corresponding to the catalytic acid/base and nucleophile residues, respectively) typical of the retaining glycanases of GH5 (45) (Fig. S11). A DALI search for structural homologs (62) yielded exclusively GH5_4 endo-glucanases and endo-xyloglucanases in the top 20 hits, with Z-scores ranging from 44 to 47 and sequence identities of 28 to 35% (Table S7).

In the complex, cellotriose was bound to both negative and positive subsites of GH5_4 MLGase. Analysis of a protein sequence alignment with 22 GH5_4 enzymes with available tertiary structural data (Fig. S11), including MLGases, xyloglucanases, and mixed-specificity β-glucanases, along with our structural observations, allowed us to identify a set of highly conserved amino acids that compose the negative subsites in this subfamily, spanning from −1 to −3. With reference to the S. copri GH5_4 structure, subsite −1 is highly conserved and comprises residues that form hydrophobic and hydrogen-bonding interactions with a glucosyl moiety: His151, His152, Asn196, His275, Try277, Trp370, and Glu380. The interactions forming the −2 subsite involve collectively Asn56, Trp370, Asn372, and Glu380. Glu380 is present in almost half of the analyzed sequences with no correlation to substrate selectivity. The −3 subsite is composed by two residues: Trp85 and His374. Of these, His374 is not conserved in the family but is also found in the Paenibacillus pabuli xyloglucanase XG5 (63) and interacts with the C6 hydroxyl group of the glucosyl residue in that position. The length of the active-site cleft, as judged from the cellotriose complex, and residue conservation suggested the possibility of a fourth negative subsite (Fig. 8).Figure 8 Active site of the GH5_4 mixed-linkage endo-glucanase.A, Segatella copri GH5_4 E331A in complex with cellotriose, in which the protein chain is represented in surface, colored according to the conservation of the residues in the subfamily GH5_4, from blue (nonconserved) to yellow (conserved). Cellotriose molecules bound to the negative (pink carbons) and positive (orange carbons) subsites are represented in sticks. B, negative subsite interactions. Amino acid labels in red indicate the catalytic nucleophile (∗) and general acid/base (∗∗) residues. C, positive subsite interactions. D, overview of the relative orientations of both cellotriose molecules in the S. copri GH5_4 E331A ternary complex. E, structural superposition of S. copri GH5_4 Glu331Ala with a Caldicellulosiruptor sp. GH5 MLGase:cellotetraose complex (PDB 5H4R, (65)) and metagenomic XEG5A:cellotriose complex (PDB 4W89, (98)) indicates the most likely positioning of the MLG chain in the positive subsites in the Michaelis (ES) complex en route to hydrolysis.

Remarkably, the cellotriose molecule in the positive subsites was bound in a highly unusual orientation, orthogonal to the active-site cleft and seemingly inconsistent with the binding of a linear MLG chain through the cleft (Fig. 8). This pose involved interactions with two tryptophan residues, Trp204 and Trp280, that stack with a single glucosyl residue and potential hydrogen bonds with the side chain of Gln278 and the main chain of Tyr277. We also note that the orientation of trisaccharide in the electron density was not unambiguous due to limited resolution in this region. To ascertain the likely trajectory of the MLG polysaccharide chain in the initial Michaelis (ES) complex, we performed structural comparisons with the metagenomic endo-xyloglucanase XEG5A in complex with cellotriose (PDB 4W89, (61)), the endo-xyloglucanase PbCel5A from Segatella bryantii in complex with xyloglucan oligosaccharides (PDB 5D9M, (64)), and the MLGase F32EG5 from Caldicellulosiruptor sp. F32 in complex with cellotetraose (PDB 5H4R, (65)) (Figs. 8E and S13). Based on this analysis, only a few residues of GH5_4 were predicted to interact with MLG. Trp204 and Trp280 likely constitute the +1 and +2 subsites, where they may stack with glucosyl residues, while Tyr277 provides an additional hydrophobic interaction in the +1 subsite. No amino acid residues were identified to form a potential +3 subsite. Together, these observations are concordant with the observed preference for cellopentaose hydrolysis through binding across the well-defined subsites −3 to +2 (Fig. S9, vide supra). The odd orientation observed for cellotriose in the positive subsites is most likely an artifact of crystallization but may hint at a potential product-release trajectory.

Unfortunately, our structural analysis does not provide a clear rationale for the exclusive selectivity of the enzyme for β(1,3) linkages in MLG (Fig. S8 cf. Fig. 1). Yet, the structural superpositions with cello-oligosaccharides (Figs. 8 and S13) clearly indicate that all-β(1,4)-glucosyl backbone regions are readily accommodated in both the negative and positive subsites. During interaction with MLG, the β(1,3)-bond spans the −1 and +1 subsites, and the marginally greater reactivity of the β(1,3)-glucosidic linkage over the β(1,4)-glucosidic linkage (66) may contribute to specificity. This specificity notably contrasts that of GH16 MLGases, which hydrolyze the β(1,4)-linkages by specifically binding β(1,3)-linkages spanning subsites −2 to −1 (43, 67). It is presently not known if the β(1,3)-linkage is accommodated in the active-site of S. copri GH5_4 by requiring the glucosyl residue in the +1 subsite to adopt a flipped orientation, as indicated by molecular dynamics simulations for a GH5_4 MLGase from the marine bacterium Zobellia galactanivorans (68). Confoundingly, structural comparison with a representative selection of MLGase, xyloglucanase, and bifunctional MLGase/xyloglucanase members of GH5_4 failed to reveal any obvious steric elements dictating the preference for MLG versus the highly branched XyG (Figs. 8 and S13).

Downstream oligosaccharide digestion

The biochemical model (Fig. 2B) implicates the GH3 member of the MLG-PUL in the production of glucose following import of the GH5_4 hydrolysis products into the periplasm by the TBDT. To explore substrate specificity, the construct GH3V23-R772, which excluded the N-terminal signal peptide and three subsequent amino acid residues, was cloned, expressed, and purified (Fig. S14). An initial screen against a panel of p-nitrophenyl glycosides revealed that GH3 was a predominant β-glucosidase, with the highest specific activity on p-nitrophenyl β-glucopyranoside and a pH optimum of 6 (Fig. S15). Specific activity on the structurally homologous p-nitrophenyl β-xylopyranoside was 5-6-fold lower due to the lack of a C-6 hydroxymethyl group. Hydrolysis of p-nitrophenyl β-galactopyranoside (C-4 epimer of PNP β-Glcp) and p-nitrophenyl β-mannopyranoside (PNP β-Manp) was not detected.

Time-dependent, qualitative analysis by HPLC indicated that the GH3 member was able to cleave diglucosides with increasing efficiency in the order gentibiose [Glc-β(1,6)-Glc], cellobiose [Glc-β(1,4)-Glc], laminaribiose [Glc-β(1,3)-Glc] (Fig. 9). Considering the downstream degradation of the products of the cell-surface GH5_4 MLGase, cellotriose, cellotetraose, and cellopentaose were each hydrolyzed by stepwise release of glucose (Fig. 9). A hexasaccharide and a heptasaccharide, obtained by partial digestion of MLG by GH5_4 (exact structures undefined, see Experimental procedures), were likewise hydrolyzed by the GH3 member (Fig. 9). Notably, in all cases, a large amount of cellobiose was observed to accumulate together with glucose, suggesting that the β-glucosidase was comparatively inefficient at cleaving the final β(1,4)-linkage of the oligosaccharide products of the GH5_4.Figure 9 Hydrolysis of oligosaccharides by GH3V23-R772. Enzyme concentration was held constant (0.008 mg/ml), as was the concentration (1 mM) of cellobiose (G4G), laminaribiose (G3G), gentibiose (G6G), cellotriose (G4G4G), cellotetraose (G4G4G4G), and cellopentaose (G4G4G4G4G). The concentrations of the MLG hexasaccharide and heptasaccharide were 180 μM and 99 μM, respectively.

As introduced above, we observed a modest but significant upregulation of a GH94 member (Fig. 1C). Although we were unable to recombinantly produce this protein in soluble form, high sequence similarity with characterized GH94 members (41) (e.g., Cellvibrio gilvus cellobiose phosphorylase, 61% sequence identity (69)) strongly suggests that it is a cellobiose phosphorylase (EC 2.4.1.20). This suggests that both glucose and cellobiose (products of the GH3 member) are transported into the cytosol, where the GH94 member mediates phosphorolysis of the latter to α-glucose-1-phosphate (Cori ester) and glucose for primary metabolism. Interestingly, a GH3 homolog from the B. ovatus (ATCC 8483) MLG-PUL also exhibited comparatively weak activity on cellobiose (43). However, the corresponding GH16 MLGase generates G4G3G and G4G4G3G as limit digest products, such that the last substrate encountered by the B. ovatus GH3 member is laminaribiose (G3G), which it hydrolyzes efficiently. Hence, the B. ovatus ATCC 8483 genome does not encode a GH94 member (http://www.cazy.org/b5380.html (41)).

Application of MLG-PULs to probe MLG metabolic potential in human metagenomes

Homologs of biochemically validated PULs serve as molecular markers of polysaccharide metabolism in metagenomes (43, 46, 56). The biochemical validation of the S. copri MLG-PUL system, including defining the strict MLG specificity of the GH5_4 member, motivated us to identify homologous MLG-PULs in other human gut bacteria in the PULDB (31, 40). As shown in Figure 2A, a variety of members of the family Prevotellaceae, including Segatella, Hoylesella, Prevotella, and Xylanibacter species, contain MLG-PULs with identical gene organization to that of S. copri DSM18205, including the vanguard GH5_4 member. This synteny also reveals the lack of conservation of the predicted GH94 glucoside phosphorylase of S. copri, which further validated its formal exclusion from the MLG-PUL, as discussed above. Exceptionally, the predicted MLG-PUL from Prevotella dentalis DSM3688 shares partial synteny with the MLG-PUL of S. copri yet is distinguished by GH94 and MFS transporter genes immediately following the GH5_4 gene, an additional GH3 gene, and relocation of the HTCS to the downstream end of the PUL (Fig. 2A).

Notably, we also observed that several Prevotellaceae instead possessed MLG-PULs that are syntenic to the characterized MLG-PUL from B. ovatus (Fig. 2A). These MLG-PULs encode a GH16_3 member (70) as the vanguard outer membrane–anchored MLGase and are found in several Bacteroides species (family Bacteroidaceae) as well as Dysgonomonadaceae (see (43) for a detailed analysis). A recent whole-genome phylogenetic analysis places Bacteroidaceae and Prevotellaceae as sister clades within a monophyletic group (71). This may suggest that the GH16_3-based MLG-PUL reflects the ancestral form, with the MLGase being replaced by a functionally similar GH5_4 MLGase in some Prevotellaceae. Otherwise, the observation of both GH5_4-based and GH16_3-based MLG-PULs in Prevotellaceae may highlight a genetic fluidity that transcends the divergence from the sister Bacteroidaceae.

We then used these MLG-PULs to interrogate the prevalence of MLG utilization across human populations, expanding our previous effort (43) here to include a broad range of Prevotellaceae and a greater diversity of samples arising from recent metagenomic studies. We used Magic-BLAST (72) to align entire short-read fecal metagenome datasets from 1977 individual humans in 13 studies to the MLG-PULs, using a stringent 95% nucleotide sequence identity threshold. Sequence coverage was subsequently calculated (Fig. S16). Across populations, MLG-PULs from known human fecal isolates of Prevotellaceae (e.g., S. copri) were the most prevalent, while MLG-PULs from oral isolates (P. dentalis, P. fusca, and P. melanogenica) were infrequently found (Fig. 10). As a negative control, the syntenic MLG-PUL of the bovine rumen isolate Xylanibacter rumincola (Fig. 2A) was never observed.Figure 10 Prevalence of Prevotellaceae MLG-PULs in humans. One thousand nine hundred seventy-seven individual human metagenomes from the indicated studies were probed by aligning reads to the PULs shown in Figure 2A, using Magic-BLAST at 95% sequence identity threshold. Individual results based including sequence coverage are shown in Fig. S16. Here, individual HGMs were considered to possess a corresponding MLG-PUL when the sequence coverage was 50% or greater. The MLG-PUL of Bacteroides ovatus, which is particularly abundant in Western European and North American HGM (43), was chosen to represent the Bacteroidaceae. The stacked bar in each dataset (rightmost) represents the percentage of individuals who carry MLG-PULs only from Prevotellaceae, only from B. ovatus, both, and none.

We also observed interesting population-based trends regarding the contribution of Prevotellaceae versus Bacteroidaceae to MLG digestion in individuals (Fig. 10). As we have shown previously using a smaller sample size (43), Bacteroidaceae-mediated MLG utilization is nearly ubiquitous in North Americans (HMP), Europeans (MetaHit), and Chinese (China 1–6), represented here by the B. ovatus MLG-PUL (see (43) for homologs). Only 20 to 30% of individuals in these cohorts carried one or more Prevotellaceae MLG-PULs and typically in combination with a Bacteroidaceae counterpart (Figs. 10 and S16). Only rarely were Prevotellaceae, the only contributors of an MLG-PUL in North American, European, and Chinese samples, most notably in Europeans. Strikingly, very few Japanese samples had either Prevotellaceae or Bacteroidaceae MLG-PULs, and where found, these were mutually exclusive. The small sample size of the Yanomami (indigenous Amazonians) precluded conclusive analysis, although the exclusive presence of Prevotellaceae was indicated in at least one individual (Fig. S16). Hadza samples (indigenous Tanzanians) were devoid of the Bacteroidaceae (B. ovatus) MLG-PUL, whereas Prevotellaceae MLG-PULs were observed in approximately half of the individuals at the 50% sequence coverage threshold (Figs. 10 and S16). Most notably, individuals from India had high prevalence of Prevotellaceae MLG-PULs, which were found in approximately two-thirds of samples. Moreover, Prevotellaceae and Bacteroidaceae MLG-PULs were essentially mutually exclusive, with very few individually carrying representatives of both families.

Overall, these observations are concordant with previous analyses based on 16S ribosomal and metagenome sequencing that indicate the segregation of individual Bacteroidaceae- and Prevotellaceae-dominated “enterotypes” between populations (11, 20). However, we note that there is clear overlap and thus functional redundancy, in individuals from postindustrial societies (N. America, Europe, and China), many of whom (ca. one-fourth) carry both Bacteroidaceae and Prevotellaceae MLG-PULs. Finally, we note that no effort was made to control for individual sample sequence quality and depth, which may contribute false negatives due to the stringent identity and coverage thresholds applied. This, coupled with the limited set of PULs used as probes, and the possibility of alternate MLG utilization systems in Bacteroidota and other phyla rationalize our inability to observe MLG-PULs in all individuals (Fig. 10, gray bars). Given the ubiquity of cereal cell walls in human diets worldwide, we would anticipate all HGM would include some form of MLG utilization system, barring a restricted diet or disease.

Conclusion

The PUL paradigm was first established for the starch utilization system of B. thetaiotaomicron (42) and has since been extended to the functional characterization of a wide range of polysaccharide- and glycan-utilization systems in diverse Bacteroides species (21, 22, 23, 24, 25, 26, 27, 28, 29, 30). Canonical PULs encode all of the proteins required to sense, capture, and completely saccharify individual, cognate substrates. Our bioinformatic, biochemical, and structural data support the proposed model (Fig. 2) for the degradation and uptake of MLG by the HGM bacteria S. copri.

This model recapitulates all key aspects of Bacteroidetes PUL systems across the cell envelope: (i) SGBP-B binds β-glucans at the cell surface, increasing the local concentration to aid hydrolysis (ii) by the specific GH5_4 MLGase. (iii) Oligosaccharide products are then captured by the MLG-specific SGBP-A/TBDT (SusC/SusD homolog) complex and shuttled into the periplasm. In light of recent advances, we note that steps 1 to 3 likely occur in a “utilisome” (51), that is, a complex of the four outer membrane–bound proteins. (iv) Further hydrolysis by the GH3 exo-β-glucosidase releases glucose, which is transported to the cytoplasm to enter primary metabolism. Cellobiose may be generated as a coproduct of this process and likewise transported to the cytosol, where it can be phosphorylyzed by an independent GH94 member. The precise specificity of the sensor domain of the HTCS is presently unknown, but analogy with the B. ovatus MLG-PUL suggests that longer oligosaccharides (not glucose) serve as inducers (39, 43).

These insights into the molecular mechanism of MLG utilization by S. copri are critical to understanding the ability of Prevotellaceae to degrade complex polysaccharides. Future studies on other PULs from S. copri and related species will help to define roles of Prevotellaceae in the human gut ecosystem, particularly in the context of dietary choices and health. Furthermore, this study provides a foundation for investigating host–microbe interactions, including targeted modulation of HGM composition in people to ameliorate metabolic or bowel diseases. For example, recent studies have demonstrated that PULs from S. copri (syn. P. copri) metagenome-assembled genomes, which are homologous to the MLG-PUL described here, are specifically enriched and upregulated in the microbiota of malnourished children during intervention with a therapeutic food formulation (73, 74). We can now unambiguously assign MLG specificity to these PULs based on the biochemistry presented here (see Fig. 4 and Table S1e in (74)). In turn, this underscores the value of plant β-glucan as a component of Microbiota-Directed Complementary Food formulations (73).

Experimental procedures

RNA sequencing

S. copri (syn. Pr. copri) DSM 18205 was obtained from the Leibniz Institute DSMZ – German Collection of Microorganisms and Cell Cultures GmbH. Bacteria were grown in an anaerobic chamber (Coy Laboratory Products Inc.) at 37 °C in an atmosphere of 90% N2, 5% CO2, and 5% H2. All media were filter sterilized and allowed to equilibrate in the chamber for at least 6 h before use. S. copri DSM18025 cultures were harvested at mid-log phase (A600nm = 0.6–0.8) from modified PYG medium (Table S8) or modified PYM medium containing 0.5 g/L barley MLG (Table S9). RNA protect (QIAGEN) was added in a 1:1 ratio to culture, followed by freezing and storage at −80 °C. Cells were pelleted and the supernatant was removed before being resuspended in PBS with 20 mg/ml lysozyme. Cells were lysed using 0.1 mm zirconium beads and a Tissue Lyser (QIAGEN). The cell lysate was centrifuged and using RNAeasy bacterial RNA extraction kit (QIAGEN), total RNA was isolated from each sample. RNA samples were assessed for quality using an Aligent BioAnalyzer and were allowed to proceed to sequencing only if the RIN score was above 8. Total RNA samples were enriched for mRNA using NEBNext rRNA Depletion Kit (bacteria). Enriched mRNA was sequenced on an Illumina MiSeq platform. Reads were assessed using FastQC (http://www.bioinformatics.babraham.ac.uk/projects/fastqc/) and MultiQC (75), trimmed with Trimmomatic (76), aligned to S. copri DSM18025 reference genome (77), and counted using Salmon (78). Differential expression of each polysaccharide growth compared to glucose grown cells was calculated using DESeq2 (79). Raw sequence reads were deposited in the Sequence Read Archive of the National Center for Biotechnology Information (NCBI) under BioProject Accession PRJNA1073563.

Gene identification, recombinant expression, and protein production and purification

S. copri DSM 18205 genomic sequences were obtained from the Integrated Microbial Genomes database of the Joint Genome Institute. Locus tags selected for recombinant expression refer to the most recent genome assembly (77), with previous locus tags (31) and GenBank accessions listed in parentheses: NQ544_01055 (Prevcop_05090, UWP52536.1; GH94), NQ544_01030 (Prevcop_05096, UWP52531.1; GH3), NQ544_01025 (Prevcop_05097, UWP52530.1; SGBP-B), NQ544_01020 (Prevcop_05098, UWP52529.1; SGBP-A), and NQ544_01010 (Prevcop_05101, UWP52527.1; GH5_4). Predicted gene products were analyzed for signal peptides using SignalP (48) and Psipred (80) and for disordered regions using structural models created by AlphaFold (81). Constructs excluding signal peptides, lipidation residues, and potential disordered N-terminal regions were amplified by PCR from S. copri CB7 (syn. P. copri DSM 18205) genomic DNA using Q5 high-fidelity polymerase (New England Biolabs) and inserted into pMCSG53 vectors through ligation-independent cloning (82); see Table S10 for primers and Figs. S1 S2, S5, and S14 for construct schematics. In a similar way, GH94M1-M831 was cloned into pET28a using the restriction enzymes NdeI and high fidelity HindIII as well as T4 DNA ligase (New England Biolabs). The GH5_4 mutants Glu197Ala (acid/base) and Glu331Ala (nucleophile) were generated using the Q5 Site-Directed Mutagenesis Kit (New England Biolabs) with GH5_4T41-E406 in pMCSG53 vector as the template DNA. All constructs were confirmed via PCR and Sanger sequencing (Genewiz).

Constructs were expressed in BL21 (DE3) cells grown in LB medium and induced with 1 mM IPTG when the optical density (A600nm) reached 0.6. During induction, cultures incubated at 16 °C with orbital shaking at 250 rpm for 48 h. Cells were harvested by centrifugation at 4367g for 30 min at 4 °C, and the resulting pellets were resuspended in lysis buffer [20 mM sodium phosphate (pH 7.4), 500 mM NaCl, 20 mM imidazole]. The cell suspensions were sonicated using a Sonic Dismembrator F55 Ultrasonic Homogenizer (Thermo Fisher Scientific) in six pulses of 30 s in 30% amplitude separated by 1 min intervals. The lysed cells were centrifuged at 4367g for 45 min.

Recombinant proteins were purified by immobilized metal ion affinity chromatography (IMAC) followed by size-exclusion chromatography on an NGC Chromatography System (Bio-Rad). After lysis of E. coli cells as described above, the soluble fraction was applied to a pre-equilibrated 5 ml HisTrap HP column (Cytiva). Washing of the column with 40 ml of binding buffer [20 mM sodium phosphate (pH 7.4), 500 mM NaCl, 20 mM imidazole] was followed by linear gradient elution buffer [20 mM sodium phosphate (pH 7.4), 500 mM NaCl, 500 mM imidazole] in 100 ml. The elution of the proteins of interest was monitored by the absorbance at 280 nm and the purity of the fractions was confirmed by SDS-PAGE. Pure fractions were pooled and concentrated using a VivaspinTM centrifugal filter (Cytiva). All protein concentrations were determined by A280 measurements, using individual calculated protein molar extinction coefficients.

For the SGBP-A construct used for crystallization experiments, an additional purification step was performed to remove the 6xHis-tag. After the first IMAC step, the fractions containing pure protein were concentrated and a buffer exchange was performed to buffer 20 mM sodium phosphate (pH 7.4) and 500 mM NaCl. Four hundred micrograms of tobacco etch virus protease were added to the concentrated protein solution overnight at 4 °C and a second IMAC was carried out to separate the protein from the cleaved 6xHis-tag. Concentrated protein samples were loaded into a Superdex 75 prep grade or Superdex 200 prep grade resin (Cytiva) size-exclusion column pre-equilibrated with a buffer containing 20 mM 4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid (Hepes) (pH 7.0) and 100 mM NaCl. Resulting fractions containing the proteins were analyzed by SDS-PAGE (Coomassie Brilliant Blue staining), combined, and concentrated using an Vivaspin 20 (Cytiva) filter.

Substrates

Barley MLG (low viscosity), curdlan, potato galactan, carob galactomannan, konjac GM, β-1,4-linked mannan, xylan, xyloglucan, yeast β-glucan were purchased from Megazyme International. Laminarin (Laminaria digitata) was obtained from Sigma Aldrich, carboxymethyl cellulose from Acros Organics, and HEC from Amresco. Xanthan gum was purchased from Spectrum, and lentinan and pustulan from Elicityl.

All oligosaccharides (cellotriose, cellotetraose, cellopentaose, laminaribiose, laminaritriose, laminaritetraose, 31-β-D-cellobiosyl-glucose (G4G3G), 32-β-D-glucosyl-cellobiose (G3G4G), 31-β-D-cellotriosyl-glucose (G4G4G3G), 32-β-D-cellobiosyl-cellobiose (G4G3G4G), 33-β-D-glucosyl-cellotriose (G3G4G4G)) were purchased from Megazyme, except cellobiose and gentibiose, which were ordered from Acros Organics and Carbosynth, respectively. Chromogenic para-nitrophenol (pNP) glycosides, including pNP α-glucoside, pNP β-glucoside, pNP α-galactoside, pNP β-galactoside, pNP β-mannoside, pNP α-xyloside, and pNP β-xyloside were purchased from Sigma Aldrich.

MLG-hexasaccharide and MLG-heptasaccharide were produced at 37 °C in a reaction containing 900 μl of 400 mM sodium citrate buffer (pH 4.5), 100 μl of ddH2O, 8 ml of 10 mg/ml barley MLG, and 1 ml of 0.002 mg/ml GH5_4C21-E406 in 40 mM sodium citrate buffer, pH 4.5, containing 0.1 mg/ml bovine serum albumin (BSA). The reaction was terminated after 10 min by heating the solution at 90 °C for 10 min. The solution was mixed with 30 ml ice-cold ethanol and was stored at –20 °C overnight. After centrifugation at 21,100g for 2 h, the supernatant was transferred to a round bottom flask and solvents were evaporated using a Rotavapor RII (BUCHI). The oligosaccharides were dissolved in 100 μl ddH2O and separated on a size-exclusion column (2.6 cm diameter, 1 m length) containing Bio-Gel P2 resin (Biorad) using a flowrate of 0.5 ml/min. Fractions were analyzed for the presence of carbohydrates using the bicinchoninic acid (BCA) reducing-sugar assay (vide infra). Selected fractions were analyzed by high-performance anionic-exchange chromatography with pulsed amperometric detection (HPAEC-PAD) using gradient I and gradient II (vide infra). Pure fractions were combined, lyophilized, and dissolved in ddH2O. The concentration of generated oligosaccharide was determined from the amount of produced glucose during a digestion of 2.5 μl oligosaccharide with 0.08 mg/ml GH3V23-R772 in a 100 μl reaction. The solution was incubated for 4 h in 40 mM NaPi (pH 6.0) at 37 °C and analyzed via HPAEC-PAD using gradient I. Glucose amounts were calculated from a glucose standard curve (5–200 μM). Solutions containing oligosaccharides were diluted to a concentration of 100 μM and analyzed via HPAEC-PAD and via an Agilent 1100/1200/1260/1290 ESI Q-TOF MS (Fig. S17).

BCA assay

Enzyme activity on polysaccharides was determined according to the BCA reducing sugar assay protocol (83) using reaction volumes of 100 μl and 100 μl of BCA solution to terminate the reaction after 6 min. Enzyme solutions were prepared to the desired concentration using buffer, also including 0.1 mg/ml BSA. After heating, the absorbance of 100 μl solution was determined at λ = 562 nm with an Epoch Microplate Spectrophotometer (Agilent BioTek). To exclude the effect of buffer or substrate, absorbance values for blanks containing only buffer and 0.1 mg/ml BSA were subtracted from the observed absorbance and amounts of generated carbohydrate were calculated using a glucose standard curve (10–150 μM). All reactions were performed in the linear range of the enzyme.

For an initial substrate screen, 10 μl of 5 mg/ml polysaccharide (barley MLG, XyG, GM, HEC, carboxymethylcellulose, xylan, curdlan, L. digitata laminarin, yeast β-glucan, and xanthan gum) or 2.5 mg/ml polysaccharide (pustulan) were incubated with 80 μl 50 mM NaPi buffer (pH 6.0) and 10 μl GH5_4C21-E406 (0.0001–0.1 mg/ml) at 37 °C. Similarly, the hydrolytic activities of GH5_4T41-E406, GH5_4T41-E406(E197A), and GH5_4T41-E406(E331A) on barley MLG were assessed by the incubation of 10 μl of polysaccharide (5 mg/ml), 10 μl enzyme (0.0001–0.2 mg/ml) at 37 °C in 40 mM NaCit buffer (pH 4.5).

The pH-rate and temperature profile of GH5_4C21-E406 were assessed in sodium citrate buffer (pH 3.0–6.0), 2-(N-morpholino)ethanesulfonic acid buffer (pH 5.5–6.0), NaPi buffer (pH 6.0–8.0), glycine sodium hydroxide buffer (Gly-OH; pH 9.0–10) and temperatures ranging from 20 to 70 °C. Reactions contained 0.5 mg/ml barley MLG and 0.01 μg/ml GH5_4C21-E406 in 40 mM buffer.

To obtain the characteristic kcat and Km values for polysaccharide degradations, reactions, containing 90 μl of polysaccharide solution (0–1.2 mg/ml), 10 μl GH5_4C21-E406 solution (0.0001–0.1 mg/ml), and a final concentration of 40 mM NaCit buffer (pH 4.5), were incubated at 50 °C. The Michaelis–Menten equation was fit to initial-rate kinetic data using OriginPro 2019 Version 9.6.0.172.

pNP glycoside assay

To assess the activity of GH3, chromogenic assay based on the hydrolysis of para-nitrophenyl glycosides was employed. Using an end-point assay set up, reactions were carried out in a volume of 100 μl reactions containing 1 mM pNP glycosides and 0.008 mg/ml GH3 at 37 °C for 10 min. The activity of GH3 was first tested against a library of pNP derivatives: pNP α-glucoside, pNP β-glucoside, pNP α-galactoside, pNP β-galactoside, pNP β-mannoside, pNP α-xyloside, and pNP β-xyloside in 40 mM sodium phosphate (pH 6.2) buffer. To monitor the pH-rate profile, GH3 was incubated with pNP β-glucoside in 40 mM sodium citrate (pH from 3.0 to 6.0; varying 0.5 between each measure), 40 mM sodium phosphate (pH from 6.0 to 8.0; varying 0.5 between each measure), and 40 mM glycine-hydroxide (pHs 8.5 and 9.5) buffers. All the reactions were terminated by the addition of 100 μl of 1 M Na2CO3 solution and monitored at 405 nm using an Epoch Microplate Spectrophotometer (Agilent BioTek). Rates were calculated using a molar extinction coefficient for p-nitrophenylate of 18,000 M−1 cm−1.

Carbohydrate analysis by HPAEC-PAD

Enzyme products were analyzed using HPAEC-PAD on a Dionex ICS-6000 system equipped with a 3 × 250 mm Dionex CarboPac PA200 IC column and a 3 × 50 mm guard column. Samples were injected via an AS-AP autosampler with a temperature-controlled sample tray maintained at 4 °C. Chromeleon 7 software was used for instrument control and initial data analysis. Chromatograms were processed subsequently using OriginPro 2019 Version 9.6.0.172 for quantitative analyses.

All samples were filtered through a Costar Spin-X centrifuge tube filter containing a 0.22 μm pore CA membrane prior to injection (10 μl). Three optimized gradients were employed for carbohydrate separation, using the solvents ddH2O (solvent A), 1.0 M NaOH (solvent B), and 1.0 M sodium acetate (NaOAc; solvent C).• Gradient I: 0 to 15 min; 10% B, 1 to 30% C linear gradient; 15 to 15.1 min; 10 to 50% B, 30 to 50% C linear gradients; 15.1 to 16 min, 50 to 10% B, 50 to 1% C exponential gradients; 16 to 20.1 min, 10% B, 1% C

• Gradient II: 0 to 20 min; 10% B, 1 to 8% C linear gradient; 20 to 20.1 min; 10 to 50% B, 8 to 50% C linear gradients; 20.1 to 21 min, 50 to10% B, 50 to1% C exponential gradients; 21 to 25.1 min, 10% B, 1% C

• Gradient III: 0 to 14 min; 10% B, 1 to 5.9% C linear gradient; 14 to 14.1 min; 10 to 50% B, 5.9 to 50% C linear gradients; 14.1 to 15 min, 50 to10% B, 50 to1% C exponential gradients; 15 to 20 min, 10% B, 1% C.

The polysaccharide digestion profile of barley MLG generated by GH5_4 at 37 °C was assessed as follows. Hydrolysis was initiated by the addition of GH5_4C21-E406 in enzyme buffer resulting in final concentrations of 0.0002 mg/ml enzyme, 8 mg/ml barley MLG, and 40 mM NaCit (pH 4.5) in 1000 μl reaction mixture. Aliquots of 100 μl were mixed with 100 μl heated NaCit buffer (40 mM, pH 4.5) after specific time intervals. The solution was heated to 90 °C for 10 min and an aliquot of 20 μl diluted to 100 μl using ddH2O. The carbohydrate mixture was separated using gradient I to monitor the production of longer and shorter oligosaccharides and gradient II for an optimized separation of shorter oligosaccharides with the same molecular weight. Standards of 100 μM oligosaccharide in 0.84 mM NaCit buffer (pH 4.5) were analyzed under identical conditions.

The limit-digestions of G4G, G4G4G, G4G4G4G, G4G4G4G4G, G3G, G3G3G, G3G3G3G, G3G4G, G4G3G, G3G4G4G, G4G3G4G, and G4G4G3G by GH5_4C21-E406 were monitored as follows. Following incubation of 10 μl GH5_4C21-E406 (0.2 mg/ml) with 90 μl of 100 μM oligosaccharide in 40 mM NaCit buffer (pH 4.5) for 30 min at 37 °C, the solution was heated at 95 °C for 10 min. After the addition of 10 μl of 1 mM D-mannitol as internal standard, carbohydrates were separated by HPAEC-PAD using gradient III.

To determine initial-rate kinetics and determine kcat and Km, hydrolysis of 90 μl oligosaccharide (1 μM–10 mM) with 10 μl of GH5_4C21-E406 at optimized concentrations was performed in 40 mM NaCit buffer (pH 4.5) at 37 °C. After 30 min, the solution was heated at 95 °C for 10 min, 10 μl of 1 mM D-mannitol were added as internal standard, and the products were separated on the column using gradient III. Cellobiose or cellotriose amounts were determined from standard curves (0.5–100 μM in 40 mM NaCit buffer (pH 4.5)). The Michaelis–Menten equation was fit to resulting initial-rate kinetic data in OriginPro 2019 Version 9.6.0.172. Free energy contributions of enzyme subsite binding to catalysis were calculated as described in (84) and references therein.

The oligosaccharide digestion profile generated by GH3 was determined by HPAEC-PAD as follows. One millimolar oligosaccharide (G4G, G4G4G, G4G4G4G, G4G4G4G4G, G3G, and G6G) was hydrolyzed by 0.008 mg/ml GH3V23-R772 in 40 mM NaPi buffer (pH 6.0) at 37 °C. After selected time intervals, aliquots of 100 μl were heated to 95 °C for 10 min, 10 μl of 1 mM D-mannitol were added as internal standard, and degradation products were separated using gradient III. Similarly, hydrolysis of MLG-oligosaccharides was performed using a final concentration of 180 μM MLG-hexasaccharide and 99 μM MLG-heptasaccharide. Due to the size of the oligosaccharides, carbohydrates were separated using gradient I.

Regiospecificity of cellopentaose hydrolysis

The regiospecificity of cellopentaose hydrolysis by GH5_4 was performed in [18O] water essentially as previously described (64, 85). 0.5 μl GH5_4C21-E406 (0.02 mg/ml), 0.5 μl 400 mM NaCit buffer (pH 4.5), and 11 μl [18O] water (97%; Cambridge Isotope Laboratories) were mixed and stored on ice. Following addition of 0.5 μl 10 mM cellopentaose, the solution was heated to 37 °C. Accounting for dilution, the final isotope abundance of [18O] water in reactions was 85%. After 10 min, the reaction was terminated by boiling the solution at 95 °C for 10 min. Controls were performed using 11 μl natural-abundance water. The isotopic distribution of product oligosaccharides was determined by mass spectrometry using a Waters 2695 Separation Module connected to a Waters Micromass ZQ Detector. The recorded data was processed and analyzed using Mnova Version 9.0.1-13254 and OriginPro 2019 Version 9.6.0.172. We confirmed that the rate of isotopic exchange between cellobiose, cellotriose, and cellopentaose and [18O] water was negligible under these conditions (85).

Affinity gel electrophoresis

To determine qualitative binding specificities of SGBP-A and SGBP-B, polysaccharides (barley MLG, XyG, GM, HEC, galactan, galactomannan, lentinan, laminarin, and xylan) were dissolved in ddH2O at 60 °C to a concentration of 10 mg/ml. Solutions containing 5 mg/ml of yeast β-glucan and β-1,4-linked mannan or 4 mg/ml curdlan were prepared by dissolving 100 mg of each polysaccharide in 10 ml 10% NaOH, followed by neutralization to pH 7 with 100% acetic acid.

Dyes and affinity gels (1 × running buffer, 1 mg/ml polysaccharide, 10% acrylamide/bisacrylamide) were prepared as described in (86) using 30% acrylamide/bisacrylamide solution (ratio 37:5.1) (Bio-Rad). Gels were loaded with 10 μl of protein solution (0.16 mg/ml protein in 2 mM Hepes, 10 mM NaCl, pH 7.0; 20% loading dye) and electrophoresis was conducted at 100 V for 180 min. The resulting protein bands were stained with Coomassie Brilliant Blue.

Isothermal titration calorimetry

ITC was performed at 25 °C using a MicroCal PEAQ-ITC (Malvern Panalytical). Protein concentrations were varied between 80 to 150 μM in 20 mM Hepes buffer, pH 7.0, containing 100 mM NaCl. Polysaccharide solutions were 2.5 to 5 mg/ml and cellotetraose was used at 1 mM. Each titration involved the injection of 2 μl carbohydrate solution, preceded by a 0.2 μl initial injection, delivered 18 times at intervals of 150 s, with stirring at 750 rpm. Binding affinities were calculated from the recorded heat changes using the MicroCal PEAQ-ITC Analysis Software Version 1.21 (Malvern Panalytical), assuming a 1:1 binding model.

Crystallization, X-ray diffraction data collection, and data processing

Crystallization conditions were screened in sitting-drop 96-well plates pipetted by a mosquito LCP robot (SPT Labtech), and subsequent optimizations were manually prepared in 24- or 48-well hanging-drop plates. All the plates were stored in room temperature. Crystals of SGBP-A in the free form were grown in 48-well plates by mixing 1 μl of SGBP-A at 25 mg/ml with 1 μl of reservoir solution [0.1 M Tris–HCl (pH 8.5) and 26% (w/v) PEG 8000], equilibrated against 200 μl solution in the reservoir. Complexes with cellopentaose were obtained by cocrystallization of SGBP-A at 31.5 mg/ml and cellotetraose at 5 mM (cellopentaose was obviously an impurity in this commercial preparation, which led to a fortuitous complex). Drops containing 5 μl of the mixture and 5 μl of the reservoir solution [0.1 M Tris–HCl (pH 8.5) and 23% (w/v) PEG 8000], with 1 ml solution in the reservoir of a 24-well plate. Crystals of SGBP-A were cryoprotected with a mixture of 70% reservoir solution and 30% ethylene glycol before flash freezing with liquid nitrogen.

GH5_4 crystallization experiments were performed in 48-well hanging drop plates. For the free structure, a sample containing 17 mg/ml of the protein and 5 mM of a mixture of the tetrasaccharides cellotetraose and MLG4 was prepared. One microliter of this mixture and one microliter of the reservoir solution (0.2 M LiCl, 19% PEG 3350) were pipetted, equilibrated against 200 μl of solution in the reservoir. The complex with cellotriose was obtained by pipetting 1 μl of a mixture of protein and oligosaccharide (80 mg/ml GH5_4T41-E406(E331A) and 10 mM cellotriose), 0.8 μl reservoir solution (2 M ammonium sulfate), and 0.2 μl xylitol, equilibrated against 200 μl ammonium sulfate. The resulting crystals were cryoprotected with 30% ethylene glycol solution before flash freezing.

X-ray diffraction data of SGBP-A in the free form and GH5_4 was collected at the 12-2 beamline from Stanford Synchrotron Radiation Lightsource (Palo Alto, EUA) equipped with a EIGER 16M PAD detector (Dectris). The crystals for SGBP-A in complex with cellopentaose and GH5_4 in complex with cellotriose were diffracted at the CMCF-ID beamline at the Canadian Light Source equipped with a Eiger X 9M detector (Dectris). Data sets were indexed, integrated, and scaled using the autoprocessing functions integrated in the beamlines, based on the XDS package (87). All the structures were determined by molecular replacement with Alphafold models (81) as search coordinates using the program Phaser (88). Generated models were refined with Phenix.refine (89), interspersed with manual cycles using the WinCoot program (90). The final structures were validated using MOLPROBITY (91). Crystallographic data are provided in Tables S2 and S7.

Analysis of MLG-PULs in human populations

Previously sequenced publicly available metagenomes of human stool samples were downloaded from NCBI in the following bioprojects: PRJNA422434, PRJEB10878, PRJEB12123, PRJEB12124, PRJEB15371, PRJEB6997, PRJDB3601, PRJNA48479, PRJEB4336, PRJNA392180, PRJNA527208, and PRJNA397112. Additional metagenomes from the Human Microbiome Project 2 study were downloaded from the Human Microbiome Bioactives Resource data portal (https://portal.microbiome-bioactives.org/). These datasets encompass population cross-sections across major continents including North American, Europe, Asia, and Africa. Detailed sequence metadata for Segatella copri, Segatella hominis, Segatella sinica, Hoylesella nanceiensis, Prevotella denticola, Prevotella fusca, Prevotella melaninogenica, Xylanibacter ruminicola, Prevotella dentalis, Bacteroides ovatus, Xylanibacter oryzae, Prevotella herbatica, Hoylesella loescheii, Segatella buccae, and Segatella paludivivens MLG-PULs (Fig. 2) are included as Supporting Information in a Tab Separated Value (TSV) file. Genome IDs, protein accessions, locus tags, etc., correspond to current annotations in GenBank. The “old_locus_tag” column in the TSV file contains locus tags which may no longer be associated with NCBI entries but have been used in previous literature. CAZyme annotations are from the PULDB (31, 40) if available and were otherwise manually assigned after inspection of sequence and structural homology.

To expedite computation, individual MLG-PUL sequences were concatenated into a single reference FASTA file. Magic-BLAST 1.7.0 (72) was used to align raw short-read metagenomic sequences to this concatenated sequence (equivalent to a “reference genome” in Magic-BLAST) on the Digital Research Alliance of Canada Advanced Research Computing cluster (https://alliancecan.ca/), analogous to previous analyses (56). A threshold of 95% identity was used to minimize false-positive alignments of reads to references. The output of Magic-BLAST was sorted, indexed, and converted from SAM format to BAM format using SAMtools version 1.13 (92). The coverage of each gene in each MLG-PUL was computed from the BAM file using the genomecov tool in BEDTools version 2.30.0 (93). Numerical analysis and visualization of the resulting coverage data was performed using the Python programming language, utilizing the pandas (https://zenodo.org/doi/10.5281/zenodo.3509134) and NumPy (94) libraries for data management. Heatmaps were rendered using the seaborn plotting library (95) with Matplotlib (96).

Data availability

S. copri DSM 18205 gene sequences were obtained from the Integrated Microbial Genomes database of the Joint Genome Institute. RNA sequencing data are available in the Sequence Read Archive (SRA) of the National Center for Biotechnology Information (NCBI) under BioProject Accession PRJNA1073563. Experimentally determined tertiary structures have been deposited in the Protein Data Bank (PDB) under accessions 9BAM, 9BAL, 9BJX, and 9BMK. All other data are presented in this article and the accompanying Supporting Information.

Supporting information

This article contains supporting information.

Conflict of interest

The authors declare that they have no conflicts of interest with the contents of this article.

Supporting information

Supplemental Tables S1–S10 and Figures S1–S17

Supplemental Data

Acknowledgments

We thank Professor Martin Hirst, Michelle Moksa, and Qi (Rachelle) Cao (Hirst Laboratory, Michael Smith Laboratories, University of British Columbia) for advice and performing RNA sequencing. We thank Dr Deepesh Panwar (Brumer Lab) for assistance in compiling Tables S8 and S9.

Author contributions

B. G., R. L. C., A. S. C. F., J. B., F. V. P., and H. B. writing–review and editing; B. G., R. L. C., A. S. C. F., and J. B. writing–original draft; B. G., R. L. C., A. S. C. F., and W. A. S. visualization; B. G., R. L. C., A. S. C. F., and J. B. methodology; B. G., R. L. C., A. S. C. F., J. B., and W. A. S. investigation; B. G., R. L. C., A. S. C. F., J. B., W. A. S., F. V. P., and H. B. formal analysis; R. L. C., A. S C. F., and J. B. data curation; A. S. C. F. and J. B. software; F. V. P. and H. B. supervision; F. V. P. and H. B. funding acquisition; H. B. conceptualization.

Funding and additional information

We thank the 10.13039/100022992 Canadian Institutes of Health Research (CIHR; Operating Grant MOP-142472, Project Grants PJ9-185659 and PJT-189972) and the 10.13039/501100000038 Natural Sciences and Engineering Research Council of Canada (NSERC; Discovery Grant RGPIN-2018-03892) for funding. B. G. was a recipient of a Four-Year Doctoral Fellowship from the University of British Columbia. J. B. was a CIHR Banting Postdoctoral Fellow (2019-2021).

Some crystallographic data was collected using Beamline 12-2 of the Stanford Synchrotron Radiation Lightsource, SLAC National Accelerator Laboratory, which is supported by the 10.13039/100000015 U.S. Department of Energy , Office of Science, Office of Basic Energy Sciences under Contract No. DE-AC02-76SF00515. The SSRL Structural Molecular Biology Program is supported by the DOE 10.13039/100006206 Office of Biological and Environmental Research , and by the 10.13039/100000002 National Institutes of Health , 10.13039/100000057 National Institute of General Medical Sciences (including P41GM103393). The contents of this publication are solely the responsibility of the authors and do not necessarily represent the official views of NIGMS or NIH. Crystallographic data was also collected on Beamline CMCF-ID at the Canadian Macromolecular Crystallography Facility of the Canadian Light Source, a national research facility of the University of Saskatchewan, which is supported by the 10.13039/501100000196 Canada Foundation for Innovation  (CFI), the 10.13039/501100000038 Natural Sciences and Engineering Research Council  (NSERC), the 10.13039/100013101 National Research Council  (NRC), the 10.13039/100022992 Canadian Institutes of Health Research  (CIHR), the 10.13039/100021824 Government of Saskatchewan , and the 10.13039/100008920 University of Saskatchewan .
==== Refs
References

1 Thursby E. Juge N. Introduction to the human gut microbiota Biochem. J. 474 2017 1823 1836 28512250
2 Cummings J.H. Fermentation in the human large-intestine - evidence and implications for health Lancet 1 1983 1206 1209 6134000
3 van de Guchte M. Blottière H.M. Doré J. Humans as holobionts: implications for prevention and therapy Microbiome 6 2018 81 29716650
4 Fan Y. Pedersen O. Gut microbiota in human metabolic health and disease Nat. Rev. Microbiol. 19 2021 55 71 32887946
5 El Kaoutari A. Armougom F. Gordon J.I. Raoult D. Henrissat B. The abundance and variety of carbohydrate-active enzymes in the human gut microbiota Nat. Rev. Microbiol. 11 2013 497 504 23748339
6 Klassen L. Xing X.H. Tingley J.P. Low K.E. King M.L. Reintjes G. Approaches to investigate selective dietary polysaccharide utilization by human gut microbiota at a functional level Front. Microbiol. 12 2021 632684
7 Bergman E.N. Energy contributions of volatile fatty-acids from the gastrointestinal-tract in various species Physiol. Rev. 70 1990 567 590 2181501
8 den Besten G. van Eunen K. Groen A.K. Venema K. Reijngoud D.J. Bakker B.M. The role of short-chain fatty acids in the interplay between diet, gut microbiota, and host energy metabolism J. Lipid Res. 54 2013 2325 2340 23821742
9 Pérez-Cobas A.E. Moya A. Gosalbes M.J. Latorre A. Colonization resistance of the gut microbiota against Clostridium difficile Antibiotics (Basel) 4 2015 337 357 27025628
10 Mancabelli L. Milani C. Lugli G.A. Turroni F. Ferrario C. van Sinderen D. Meta-analysis of the human gut microbiome from urbanized and pre-agricultural populations Environ. Microbiol. 19 2017 1379 1390 28198087
11 Arumugam M. Raes J. Pelletier E. Le Paslier D. Yamada T. Mende D.R. Enterotypes of the human gut microbiome Nature 473 2011 174 180 21508958
12 Huttenhower C. Gevers D. Knight R. Abubucker S. Badger J.H. Chinwalla A.T. Structure, function and diversity of the healthy human microbiome Nature 486 2012 207 214 22699609
13 De Filippis F. Pasolli E. Tett A. Tarallo S. Naccarati A. De Angelis M. Distinct genetic and functional traits of human intestinal Prevotella copri strains are associated with different habitual diets Cell Host Microbe 25 2019 444 453 30799264
14 Trefflich I. Jabakhanji A. Menzel J. Blaut M. Michalsen A. Lampen A. Is a vegan or a vegetarian diet associated with the microbiota composition in the gut? Results of a new cross-sectional study and systematic review Crit. Rev. Food Sci. Nutr. 60 2020 2990 3004 31631671
15 Hitch T.C.A. Bisdorf K. Afrizal A. Riedel T. Overmann J. Strowig T. A taxonomic note on the genus Prevotella: description of four novel genera and emended description of the genera Hallella and Xylanibacter Syst. Appl. Microbiol. 45 2022 126354
16 Blanco-Míguez A. Gálvez E.J.C. Pasolli E. De Filippis F. Amend L. Huang K.D. Extension of the Segatella copri complex to 13 species with distinct large extrachromosomal elements and associations with host conditions Cell Host Microbe 31 2023 1804 1819 37883976
17 Gellman R.H. Olm M.R. Terrapon N. Enam F. Higginbottom S.K. Sonnenburg J.L. Hadza Prevotella require diet-derived microbiota-accessible carbohydrates to persist in mice Cell Rep. 42 2023 113233
18 Vangay P. Johnson A.J. Ward T.L. Al-Ghalith G.A. Shields-Cutler R.R. Hillmann B.M. US Immigration westernizes the human gut microbiome Cell 175 2018 962 972 30388453
19 Purushe J. Fouts D.E. Morrison M. White B.A. Mackie R.I. Coutinho P.M. Comparative genome analysis of Prevotella ruminicola and Prevotella bryantii: insights into their environmental niche Microb. Ecol. 60 2010 721 729 20585943
20 Yeoh Y.K. Sun Y. Ip L.Y.T. Wang L. Chan F.K.L. Miao Y.L. Prevotella species in the human gut is primarily comprised of Prevotella copri, Prevotella stercorea and related lineages Sci. Rep. 12 2022 9055 35641510
21 Grondin J.M. Tamura K. Déjean G. Abbott D.W. Brumer H. Polysaccharide utilization loci: fueling microbial communities J. Bacteriol. 199 2017 e00860-16
22 Briggs J.A. Grondin J.M. Brumer H. Communal living: glycan utilization by the human gut microbiota Environ. Microbiol. 23 2021 15 35 33185970
23 Ndeh D. Gilbert H.J. Biochemistry of complex glycan depolymerisation by the human gut microbiota Fems Microbiol. Rev. 42 2018 146 164 29325042
24 Tamura K. Brumer H. Glycan utilization systems in the human gut microbiota: a gold mine for structural discoveries Curr. Opin. Struct. Biol. 68 2021 26 40 33285501
25 Munoz-Munoz J. Ndeh D. Fernandez-Julia P. Walton G. Henrissat B. Gilbert H.J. Sulfation of arabinogalactan proteins confers privileged nutrient status to Bacteroides plebeius mBio 12 2021 e0136821
26 Ostrowski M.P. La Rosa S.L. Kunath B.J. Robertson A. Pereira G. Hagen L.H. Mechanistic insights into consumption of the food additive xanthan gum by the human gut microbiota Nat. Microbiol. 7 2022 556 569 35365790
27 Robb C.S. Hobbs J.K. Pluvinage B. Reintjes G. Klassen L. Monteith S. Metabolism of a hybrid algal galactan by members of the human gut microbiome Nat. Chem. Biol. 18 2022 501 510 35289327
28 Crouch L.I. Urbanowicz P.A. Basle A. Cai Z.P. Liu L. Voglmeir J. Plant N-glycan breakdown by human gut Bacteroides Proc. Natl. Acad. Sci. U. S. A. 119 2022 e2208168119
29 Rønne M.E. Tandrup T. Madsen M. Hunt C.J. Myers P.N. Moll J.M. Three alginate lyases provide a new gut Bacteroides ovatus isolate with the ability to grow on alginate Appl. Environ. Microbiol. 89 2023 1 25
30 Brown H.A. DeVeaux A.L. Juliano B.R. Photenhauer A.L. Boulinguiez M. Bornschein R.E. BoGH13ASus from Bacteroides ovatus represents a novel α-amylase used for Bacteroides starch breakdown in the human gut Cell Mol. Life Sci. 80 2023 232 37500984
31 Terrapon N. Lombard V. Drula E. Lapébie P. Al-Masaudi S. Gilbert H.J. PULDB: the expanded database of polysaccharide utilization loci Nucleic Acids Res. 46 2018 D677 D683 29088389
32 Fehlner-Peach H. Magnabosco C. Raghavan V. Scher J.U. Tett A. Cox L.M. Distinct polysaccharide utilization profiles of human intestinal Prevotella copri isolates Cell Host Microbe 26 2019 680 690 31726030
33 Rosewarne C.P. Pope P.B. Cheung J.L. Morrison M. Analysis of the bovine rumen microbiome reveals a diversity of Sus-like polysaccharide utilization loci from the bacterial phylum Bacteroidetes J. Ind. Microbiol. Biotechnol. 41 2014 601 606 24448980
34 Linares-Pastén J.A. Hero J.S. Pisa J.H. Teixeira C. Nyman M. Adlercreutz P. Novel xylan-degrading enzymes from polysaccharide utilizing loci of Prevotella copri DSM18205 Glycobiology 31 2021 1330 1349 34142143
35 Hayashi H. Shibata K. Sakamoto M. Tomita S. Benno Y. Prevotella copri sp nov and Prevotella stercorea sp nov., isolated from human faeces Int. J. Syst. Evol. Microbiol. 57 2007 941 946 17473237
36 Collins H.M. Burton R.A. Topping D.L. Liao M.L. Bacic A. Fincher G.B. Variability in fine structures of noncellulosic cell wall polysaccharides from cereal grains: potential importance in human health and nutrition Cereal Chem. 87 2010 272 282
37 El Khoury D. Cuda C. Luhovyy B.L. Anderson G.H. Beta glucan: health benefits in obesity and metabolic syndrome J. Nutr. Metab. 2012 2012 851362 22187640
38 Othman R.A. Moghadasian M.H. Jones P.J.H. Cholesterol-lowering effects of oat β-glucan Nutr. Rev. 69 2011 299 309 21631511
39 Martens E.C. Lowe E.C. Chiang H. Pudlo N.A. Wu M. McNulty N.P. Recognition and degradation of plant cell wall polysaccharides by two human gut symbionts Plos Biol. 9 2011 e1001221
40 Terrapon N. Lombard V. Gilbert H.J. Henrissat B. Automatic prediction of polysaccharide utilization loci in Bacteroidetes species Bioinformatics 31 2015 647 655 25355788
41 Drula E. Garron M.L. Dogan S. Lombard V. Henrissat B. Terrapon N. The carbohydrate-active enzyme database: functions and literature Nucleic Acids Res. 50 2022 D571 D577 34850161
42 Foley M.H. Cockburn D.W. Koropatkin N.M. The Sus operon: a model system for starch uptake by the human gut Bacteroidetes Cell Mol. Life Sci. 73 2016 2603 2617 27137179
43 Tamura K. Hemsworth G.R. DéJean G. Rogers T.E. Pudlo N.A. Urs K. Molecular mechanism by which prominent human gut Bacteroidetes utilize mixed-linkage beta-glucans, major health-promoting cereal polysaccharides Cell Rep. 21 2017 417 430 29020628
44 Attia M.A. Brumer H. Recent structural insights into the enzymology of the ubiquitous plant cell wall glycan xyloglucan Curr. Opin. Struct. Biol. 40 2016 43 53 27475238
45 Aspeborg H. Coutinho P.M. Wang Y. Brumer H. Henrissat B. Evolution, substrate specificity and subfamily classification of glycoside hydrolase family 5 (GH5) BMC Evol. Biol. 12 2012 186 22992189
46 Larsbrink J. Rogers T.E. Hemsworth G.R. McKee L.S. Tauzin A.S. Spadiut O. A discrete genetic locus confers xyloglucan metabolism in select human gut Bacteroidetes Nature 506 2014 498 502 24463512
47 Grondin J.M. Déjean G. Van Petegem F. Brumer H. Cell surface xyloglucan recognition and hydrolysis by the human gut commensal Bacteroides uniformis Appl. Environ. Microbiol. 88 2022 1 14
48 Teufel F. Armenteros J.J.A. Johansen A.R. Gíslason M.H. Pihl S.I. Tsirigos K.D. SignalP 6.0 predicts all five types of signal peptides using protein language models Nat. Biotechnol. 40 2022 1023 1025 34980915
49 Dalbey R.E. Wang P. van Dijl J.M. Membrane proteases in the bacterial protein secretion and quality control pathway Microbiol. Mol. Biol. Rev. 76 2012 311 330 22688815
50 Pollet R.M. Martin L.M. Koropatkin N.M. TonB-dependent transporters in the Bacteroidetes: unique domain structures and potential functions Mol. Microbiol. 115 2021 490 501 33448497
51 White J.B.R. Silale A. Feasey M. Heunis T. Zhu Y.L. Zheng H. Outer membrane utilisomes mediate glycan uptake in gut Bacteroidetes Nature 618 2023 583 589 37286596
52 Tamura K. Foley M.H. Gardill B.R. Dejean G. Schnizlein M. Bahr C.M.E. Surface glycan-binding proteins are essential for cereal beta-glucan utilization by the human gut symbiont Bacteroides ovatus Cell Mol. Life Sci. 76 2019 4319 4340 31062073
53 Tauzin A.S. Kwiatkowski K.J. Orlovsky N.I. Smith C.J. Creagh A.L. Haynes C.A. Molecular dissection of xyloglucan recognition in a prominent human gut symbiont mBio 7 2016 e02134-15
54 Koropatkin N.M. Martens E.C. Gordon J.I. Smith T.J. Starch catabolism by a prominent human gut symbiont is directed by the recognition of amylose helices Structure 16 2008 1105 1115 18611383
55 Bolam D.N. Koropatkin N.M. Glycan recognition by the Bacteroidetes Sus-like systems Curr. Opin. Struct. Biol. 22 2012 563 569 22819666
56 Déjean G. Tamura K. Cabrera A. Jain N. Pudlo N.A. Pereira G. Synergy between cell surface glycosidases and glycan-binding proteins dictates the utilization of specific beta(1,3)-glucans by human gut Bacteroides mBio 11 2020 e00095-20
57 Arnal G. Stogios P.J. Asohan J. Attia M.A. Skarina T. Viborg A.H. Substrate specificity, regiospecificity, and processivity in glycoside hydrolase family 74 J. Biol. Chem. 294 2019 13233 13247 31324716
58 Hrmova M. Schwerdt J.G. Molecular mechanisms of processive glycoside hydrolases underline catalytic pragmatism Biochem. Soc. Trans. 51 2023 1387 1403 37265403
59 Davies G.J. Wilson K.S. Henrissat B. Nomenclature for sugar-binding subsites in glycosyl hydrolases Biochem. J. 321 1997 557 559 9020895
60 McGregor N. Arnal G. Brumer H. Quantitative kinetic characterization of glycoside hydrolases using high-performance anion-exchange chromatography (HPAEC) Abbott D.W. Lammerts van Bueren A. Protein-Carbohydrate Interactions: Methods and Protocols 2017 Springer New York, NY 15 25
61 Krissinel E. Henrick K. Inference of macromolecular assemblies from crystalline state J. Mol. Biol. 372 2007 774 797 17681537
62 Holm L. Kääriäinen S. Rosenström P. Schenkel A. Searching protein structure databases with DaliLite v.3 Bioinformatics 24 2008 2780 2781 18818215
63 Gloster T.M. Ibatullin F.M. Macauley K. Eklöf J.M. Roberts S. Turkenburg J.P. Characterization and three-dimensional structures of two distinct bacterial xyloglucanases from families GH5 and GH12 J. Biol. Chem. 282 2007 19177 19189 17376777
64 McGregor N. Morar M. Fenger T.H. Stogios P. Lenfant N. Yin V. Structure-function analysis of a mixed-linkage β-glucanase/xyloglucanase from the key ruminal Bacteroidetes Prevotella bryantii B14 J. Biol. Chem. 291 2016 1175 1197 26507654
65 Meng D.D. Liu X. Dong S. Wang Y.F. Ma X.Q. Zhou H.X. Structural insights into the substrate specificity of a glycoside hydrolase family 5 lichenase from Caldicellulosiruptor sp F32 Biochem. J. 474 2017 3373 3389 28838949
66 Capon B. Mechanism in carbohydrate chemistry Chem. Rev. 69 1969 407 498
67 Planas A. Bacterial 1,3-1,4-β-glucanases:: structure, function and protein engineering Biochim. Biophys. Acta 1543 2000 361 382 11150614
68 Dorival J. Ruppert S. Gunnoo M. Orlowski A. Chapelais-Baron M. Dabin J. The laterally acquired GH5 ZgEngAGH5_4 from the marine bacterium Zobellia galactanivorans is dedicated to hemicellulose hydrolysis Biochem. J. 475 2018 3609 3628 30341165
69 Hidaka M. Kitaoka M. Hayashi K. Wakagi T. Shoun H. Fushinobu S. Structural dissection of the reaction mechanism of cellobiose phosphorylase Biochem. J. 398 2006 37 43 16646954
70 Viborg A.H. Terrapon N. Lombard V. Michel G. Czjzek M. Henrissat B. A subfamily roadmap of the evolutionarily diverse glycoside hydrolase family 16 (GH16) J. Biol. Chem. 294 2019 15973 15986 31501245
71 García-López M. Meier-Kolthoff J.P. Tindall B.J. Gronow S. Woyke T. Kyrpides N.C. Analysis of 1,000 type-strain genomes improves taxonomic classification of Bacteroidetes Front Microbiol. 10 2019 2083 31608019
72 Boratyn G.M. Thierry-Mieg J. Thierry-Mieg D. Busby B. Madden T.L. Magic-BLAST, an accurate RNA-seq aligner for long and short reads BMC Bioinformatics 20 2019 405 31345161
73 Hibberd M.C. Webber D.M. Rodionov D.A. Henrissat S. Chen R.Y. Zhou C. Bioactive glycans in a microbiome-directed food for children with malnutrition Nature 625 2023 157 165 38093016
74 Chang H.W. Lee E.M. Wang Y. Zhou C.Y.S. Pruss K.M. Henrissat S. Prevotella copri and microbiota members mediate the beneficial effects of a therapeutic food for malnutrition Nat. Microbiol. 9 2024 922 937 38503977
75 Ewels P. Magnusson M. Lundin S. Käller M. MultiQC: summarize analysis results for multiple tools and samples in a single report Bioinformatics 32 2016 3047 3048 27312411
76 Bolger A.M. Lohse M. Usadel B. Trimmomatic: a flexible trimmer for illumina sequence data Bioinformatics 30 2014 2114 2120 24695404
77 Cheng A.G. Ho P.Y. Aranda-Díaz A. Jain S. Yu F.Q.B. Meng X.D. Design, construction, and in vivo augmentation of a complex gut microbiome Cell 185 2022 3617 3636 36070752
78 Patro R. Duggal G. Love M.I. Irizarry R.A. Kingsford C. Salmon provides fast and bias-aware quantification of transcript expression Nat. Methods 14 2017 417 419 28263959
79 Love M.I. Huber W. Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2 Genome Biol. 15 2014 550 25516281
80 Buchan D.W.A. Jones D.T. The PSIPRED protein analysis workbench: 20 years on Nucleic Acids Res. 47 2019 W402 W407 31251384
81 Jumper J. Evans R. Pritzel A. Green T. Figurnov M. Ronneberger O. Highly accurate protein structure prediction with AlphaFold Nature 596 2021 583 589 34265844
82 Stols L. Gu M.Y. Dieckman L. Raffen R. Collart F.R. Donnelly M.I. A new vector for high-throughput, ligation-independent cloning encoding a tobacco etch virus protease cleavage site Protein Expr. Purif. 25 2002 8 15 12071693
83 Arnal G. Attia M.A. Asohan J. Lei Z. Golisch B. Brumer H. A low-volume, parallel copper-bicinchoninic acid (BCA) assay for glycoside hydrolases Methods Mol. Biol. 2657 2023 3 14 37149519
84 Saura-Valls M. Fauré R. Brumer H. Teeri T.T. Cottaz S. Driguez H. Active-site mapping of a Populus xyloglucan endo-transglycosylase with a library of xylogluco-oligosaccharides J. Biol. Chem. 283 2008 21853 21863 18511421
85 Schagerlöf H. Nilsson C. Gorton L. Tjerneld F. Stålbrand H. Cohen A. Use of 18O water and ESI-MS detection in subsite characterisation and investigation of the hydrolytic action of an endoglucanase Anal. Bioanal. Chem. 394 2009 1977 1984 19543714
86 Cockburn D. Wilkens C. Svensson B. Affinity electrophoresis for analysis of catalytic module-carbohydrate interactions Methods Mol. Biol. 1588 2017 119 127 28417364
87 Kabsch W. XDS Acta Crystallogr. D Biol. Crystallogr. 66 2010 125 132 20124692
88 McCoy A.J. Grosse-Kunstleve R.W. Adams P.D. Winn M.D. Storoni L.C. Read R.J. Phaser crystallographic software J. Appl. Crystallogr. 40 2007 658 674 19461840
89 Afonine P.V. Grosse-Kunstleve R.W. Echols N. Headd J.J. Moriarty N.W. Mustyakimov M. Towards automated crystallographic structure refinement with phenix.refine Acta Crystallogr. D Biol. Crystallogr. 68 2012 352 367 22505256
90 Emsley P. Cowtan K. Coot: model-building tools for molecular graphics Acta Crystallogr. D Biol. Crystallogr. 60 2004 2126 2132 15572765
91 Williams C.J. Headd J.J. Moriarty N.W. Prisant M.G. Videau L.L. Deis L.N. MolProbity: more and better reference data for improved all-atom structure validation Protein. Sci. 27 2018 293 315 29067766
92 Li H. Handsaker B. Wysoker A. Fennell T. Ruan J. Homer N. The sequence alignment/map format and SAMtools Bioinformatics 25 2009 2078 2079 19505943
93 Quinlan A.R. Hall I.M. BEDTools: a flexible suite of utilities for comparing genomic features Bioinformatics 26 2010 841 842 20110278
94 Harris C.R. Millman K.J. van der Walt S.J. Gommers R. Virtanen P. Cournapeau D. Array programming with NumPy Nature 585 2020 357 362 32939066
95 Waskom M.L. seaborn: statistical data visualization J. Open Source Softw. 6 2021 3021
96 Hunter J.D. Matplotlib: a 2D graphics environment Comput. Sci. Eng. 9 2007 90 95
97 Parte A.C. Carbasse J.S. Meier-Kolthoff J.P. Reimer L.C. Göker M. List of Prokaryotic names with standing in nomenclature (LPSN) moves to the DSMZ Int. J. Syst. Evol. Microbiol. 70 2020 5607 5612 32701423
98 dos Santos C.R. Cordeiro R.L. Wong D.W.S. Murakami M.T. Structural basis for xyloglucan specificity and α-d-Xylp(1 → 6)-D-Glcp recognition at the -1 subsite within the GH5 family Biochemistry 54 2015 1930 1942 25714929
