==== Front mSystems mSystems mSystems mSystems 2379-5077 American Society for Microbiology 1752 N St., N.W., Washington, DC 37289197 00014-23 10.1128/msystems.00014-23 msystems.00014-23 Research Article open-peer-reviewOpen Peer Reviewenvironmental-microbiologyEnvironmental MicrobiologyCommunity- and genome-based evidence for a shaping influence of redox potential on bacterial protein evolution https://orcid.org/0000-0002-0687-5890 Dick Jeffrey M. 1 Conceptualization Data curation Formal analysis Investigation Writing – original draft jeff@chnosz.net Meng Delong 2 Funding acquisition Investigation Writing – review and editing delong.meng@csu.edu.cn 1 Key Laboratory of Metallogenic Prediction of Nonferrous Metals and Geological Environment Monitoring of Ministry of Education, School of Geosciences and Info-Physics, Central South University , Changsha, China 2 Key Laboratory of Biometallurgy of Ministry of Education, School of Minerals Processing and Bioengineering, Central South University , Changsha, China Editor Schadt Christopher W. Oak Ridge National Laboratory , USA Ad hoc peer reviewer Flamholz Avi I. California Institute of Technology , Pasadena, California, USA Address correspondence to Jeffrey M. Dick, jeff@chnosz.net Address correspondence to Delong Meng, delong.meng@csu.edu.cn The authors declare no conflict of interest. 08 6 2023 May-Jun 2023 08 6 2023 8 3 e00014-2307 1 2023 28 2 2023 Copyright © 2023 Dick and Meng. 2023 Dick and Meng https://creativecommons.org/licenses/by/4.0/ This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International license. ABSTRACT Despite deep interest in how environments shape microbial communities, whether redox conditions influence the sequence composition of genomes is not well known. We predicted that the carbon oxidation state (ZC) of protein sequences would be positively correlated with redox potential (Eh). To test this prediction, we used taxonomic classifications for 68 publicly available 16S rRNA gene sequence data sets to estimate the abundances of archaeal and bacterial genomes in river & seawater, lake & pond, geothermal, hyperalkaline, groundwater, sediment, and soil environments. Locally, ZC of community reference proteomes (i.e., all the protein sequences in each genome, weighted by taxonomic abundances but not by protein abundances) is positively correlated with Eh corrected to pH 7 (Eh7) for the majority of data sets for bacterial communities in each type of environment, and global-scale correlations are positive for bacterial communities in all environments. In contrast, archaeal communities show approximately equal frequencies of positive and negative correlations in individual data sets, and a positive pan-environmental correlation for archaea only emerges after limiting the analysis to samples with reported oxygen concentrations. These results provide empirical evidence that geochemistry modulates genome evolution and may have distinct effects on bacteria and archaea. IMPORTANCE The identification of environmental factors that influence the elemental composition of proteins has implications for understanding microbial evolution and biogeography. Millions of years of genome evolution may provide a route for protein sequences to attain incomplete equilibrium with their chemical environment. We developed new tests of this chemical adaptation hypothesis by analyzing trends of the carbon oxidation state of community reference proteomes for microbial communities in local- and global-scale redox gradients. The results provide evidence for widespread environmental shaping of the elemental composition of protein sequences at the community level and establish a rationale for using thermodynamic models as a window into geochemical effects on microbial community assembly and evolution. KEYWORDS protein evolution oxidation state redox potential Eh–pH diagram global analysis geochemistry thermodynamics Key Research and Development Program of Hunan Province 2020WK2022, 2022SK2076 Meng Delong Natural Science Foundation of Changsha kq2202089 Dick Jeffrey M. cover-dateMay/June 2023 ==== Body pmcElemental stoichiometry of biomolecules affects elemental fluxes and trophic interactions in microbial ecosystems. However, relatively little is known about how environments may shape the elemental composition of protein sequences on evolutionary timescales. A deeper understanding of the relationship between environmental conditions and protein sequence evolution would complement efforts to develop geochemical proxies and molecular records of past environments (1, 2) and could provide new insight into the chemical factors that influence selection for taxa in microbial communities. Evolutionary processes have had many millions of years to potentially be shaped by geological conditions. For example, within various genomes of opisthokonts (a major group of eukaryotes), the carbon oxidation states (Z C) of proteins with gene ages assigned by phylostratigraphy—a technique in which conservation levels of orthologous genes among species are used to estimate ages of gene families—increase after the origin of cellular organisms at ca. 4.29 Ga (billion years ago), but the trends become more diverse as the lineages diverge after 1.1 Ga (3). The signal of oxidation derived from the elemental composition of protein sequences is consistent with the rise of atmospheric oxygen on Earth. A linkage between genome evolution and environmental conditions is conveyed by the hypothesis that, over evolutionary time, the elemental composition of genomically coded protein sequences may approach incomplete metastable equilibrium (a type of energy minimization) for given environmental conditions. Data for microbial communities enable new tests of this chemical adaptation hypothesis. Genomic analyses have revealed intriguing patterns of elemental usage in protein sequences from different species (see reference [4] and references therein). However, observations for communities rather than species can be compared more directly with environmental measurements. Therefore, in this study, we analyzed community reference proteomes, which take into account the sequences of protein-coding genes from genomes available in the National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database, as well as the abundances of community members inferred from high-throughput 16S rRNA gene sequences. Actual protein abundances are not available in RefSeq, so reference proteomes for species obtained by summing the amino acid compositions of all coding sequences in the genome are based on the assumption that each sequence is equally important. Reference proteomes for species can be aggregated to higher taxonomic levels and then multiplied by taxonomic abundances to obtain community reference proteomes. Although metatranscriptomic and/or metaproteomic data would be desirable to obtain a more accurate picture of the biomolecular composition of living communities, we chose to analyze community reference proteomes for two main reasons. First, we aimed to test the hypothesis of chemical adaptation of proteins at evolutionary timescales. Reference proteomes, or potentially expressed proteins, depend only on genome sequences, but metaproteomes, or actually expressed proteins, depend on genome sequences as well as on cellular physiology. By using reference proteomes, the analysis is focused on differences that emerge over evolutionary rather than physiological timescales. Second, data sets with redox potential measurements and 16S rRNA gene sequences are much more widely available than those for metagenomes, metatranscriptomes, or metaproteomes. To illustrate this semi-quantitatively, searches performed on Google Scholar on 26 December 2022 with the queries “16S rRNA” “redox potential”, “metagenome” “redox potential”, “metatranscriptome” “redox potential”, and “metaproteome” “redox potential” gave about 15,200, 2,940, 521, and 212 results, respectively. Published 16S rRNA-based studies, therefore, provide the most extensive source of data that can be used to compare microbial communities with environmental measurements. Recently, comparisons of this type were performed for data sets for hydrothermal systems, stratified water bodies, and shale gas wells (5). That study identified multiple field sites where carbon oxidation state of community reference proteomes was aligned with measured oxygen concentrations; moreover, the 16S rRNA-based estimates show good correspondence with Z C of proteins inferred from shotgun metagenomes (5). However, to date, there has been no quantitative comparison of the carbon oxidation states of protein sequences with redox potential measurements at a global scale. Concentrations of dissolved oxygen (O2) diminish to levels that are below detection in many sediments (e.g., within several millimeters to centimeters of the sediment–water interface (SWI); see reference [6]) and other reducing environments. In anoxic settings, measurements of hydrogen (H2) concentration can be used to monitor redox conditions (7). Oxygen and hydrogen, and other relevant oxidants and reductants, do not attain equilibrium in most environments. Nevertheless, a single redox scale is needed for a global comparison of microbial habitats. Furthermore, the redox property should be detectable in all environments; neither O2 nor H2 concentration satisfies this criterion, but electrode measurements do. For this reason, a measurement referred to as oxidation–reduction potential or redox potential (Eh) was selected in this study. Redox potential denotes the tendency for oxidation–reduction reactions to occur, with higher (more positive) values indicating a greater propensity for loss of electrons. The Eh scale is expressed as an electrical potential (i.e., voltage) with reference to the standard hydrogen electrode (SHE). Field measurements typically use a redox probe consisting of a platinum indicator electrode and an internal Ag/AgCl reference electrode connected to a potentiometer. Redox probes must be periodically calibrated with solutions of known redox potential (e.g., ZoBell’s solution), and a temperature-dependent conversion is used to convert the meter readings from volts versus Ag/AgCl to volts versus SHE (8). A rigorous chemical interpretation of Eh measurements in natural systems is challenging because they represent mixed potentials that are affected by the presence of various, often not well-defined, electroactive species that are generally not in mutual equilibrium (9). Despite its chemical complexity, Eh appears in many environmental microbiology studies, calling for new interdisciplinary approaches to examine the ecological relevance of this parameter. The specific aims of this study are first, to develop a multiscale (local to global) perspective on differences of carbon oxidation states of community reference proteomes in relation to redox potential; second, to characterize these differences in phylogenetic domains (bacteria and archaea); and finally to examine practical issues such as whether Eh or O2 concentration is a better predictor of carbon oxidation state, the robustness of the results to different primer sets for the 16S rRNA gene, and comparison with metaproteomic data. The results provide a novel hypothesis-driven picture of global microbial communities as chemical entities and develop the rationale for integrating geochemical thermodynamics into evolutionary models for microbial ecosystems. MATERIALS AND METHODS Thermodynamic calculations We used group additivity parameters that include pH-dependent ionization of sidechain and terminal groups (10) as implemented in the CHNOSZ package (11) to calculate standard Gibbs energies from amino acid compositions of reference proteomes. Chemical formulas of reference proteomes were normalized by the total number of amino acids, thus eliminating protein length as a variable in the calculations. The chemical system was defined using a minimum number of thermodynamic components, also known as basis species. We modified the previously described QEC basis species (glutamine, glutamic acid, cysteine, H2O, and O2) (12) by substituting H+ and e − for O2; the addition of electronic charge is needed for calculating an Eh–pH diagram. The chemical activities (a) of some of the basis species were assigned constant values (loga Gln = −3.2, loga Glu = −4.5, loga Cys = −3.6, and loga H2O = 0), leaving Eh and pH as variables for making a relative stability diagram. The relative stability fields on this diagram represent the reference proteome with the lowest Gibbs energy of formation from the basis species (i.e., the least unstable reference proteome) compared to the others. Data sources We used literature searches to find field-based data for different environments and field- and laboratory-based data (including mesocosms) for soils and sediments; experiments described as bioleaching were not included. All published data sets that we could find as of July 2022 were included in our compilation if (i) there were sufficient metadata to match sample names to database accession numbers, (ii) demultiplexed 16S rRNA gene sequences were available in the NCBI Sequence Read Archive (SRA) (13), and (iii) there were at least 20 biological samples with an associated Eh range of at least 100 mV before correction to pH 7. The sample number cutoff was reduced to 15, 10, and 8 for groundwater, geothermal, and hyperalkaline environments, respectively, to compile enough data sets for global comparison. Data set names used here, NCBI BioProject accession numbers, and references are listed in Table 1. Each study (i.e., source publication) corresponds to one data set, with the following exceptions: two data sets were compiled from each of the studies of Ghannam et al. (14) (Port Microbes, the two data sets are for samples collected on a prefilter or postfilter), and Power et al. (15) (New Zealand Hot Springs, the two data sets are for acidic or circumneutral to alkaline samples). Data for Winogradsky columns (16, 17) were analyzed separately from the main compilation. TABLE 1 Data sets analyzed in this study a River & seawater Pearl River Estuary (PRJNA319446 and PRJNA357334 [18]), Sansha Yongle Blue Hole (PRJNA503500 [19, 20]), Port Microbes–prefilter and postfilter (PRJNA542685 and PRJNA542890 [14]), Bahe River (PRJNA588356 [21]), Nu River (PRJNA663210 [22]), Maozhou River (PRJNA681688 [23]), Plastisphere (PRJNA717904 [24]), Three Gorges Reservoir (PRJNA733826 [25]), and Taxco AMD (PRJNA801253 [26]) Lake & pond Microbial Mat, Kiritimati Atoll (PRJNA174394 [27]), Xidong Reservoir, Xiamen (PRJNA315049 [28]), Ursu Lake, Romania (PRJNA395513 [29]), Xiamen Reservoirs and Ponds (PRJNA407260 [30]), Keweenaw Waterway, Michigan (PRJNA489447 [31]), German Kettle Holes (PRJNA641761 [32]), Lake Kinneret, Israel (PRJEB39923 [33]), Monegros Desert, Spain (PRJNA429605 [34]), and Lake Varese, Italy (PRJNA694444 [35]) Geothermal areas New Zealand Hot Springs–acidic and circumneutral to alkaline (PRJEB24353 [15]), Eastern Tibetan Plateau (PRJNA592622 [36]), Uzon Caldera (PRJNA623081 [37]), and Southern Tibetan Plateau (PRJNA638734 [38]) Hyperalkaline fluids CROMO 1 (PRJNA289273 [39]), Samail Ophiolite (PRJNA352492 [40]), Santa Elena Ophiolite (PRJNA361138 [41]), Voltri Massif (PRJNA685937 [42]), CROMO 2 (PRJNA690585 [43]), and Samail Ophiolite Packers (PRJNA743134 [44]) Groundwater Sarnia nZVI Injection (PRJNA308958 [45]), Hetao Basin (PRJNA350383 [46]), Jianghan Plain (PRJNA377933 [47]), Ohio Aquifers (PRJNA387583 [48]), Rayong Province (PRJNA434769 [49]), Mezquital Valley (PRJNA488796 [50]), Hainich Critical Zone (PRJEB33032 [51]), Po Plain (PRJNA667833 [52]), New Zealand Aquifers (PRJNA699054 [53]), and Aquifer SE of Melbourne (PRJNA861729 [54]) Sediment Mai Po Wetland (PRJEB12429 and PRJEB12432 [55]), Baltic Sea Swedish Coast (PRJNA322450 [56]), Finnish Boreal Lakes (PRJNA349972 [57]), Honghu Lake (PRJNA352457 [58]), Marine and Freshwater Sediments (PRJNA393823 [59]), Daya Bay (PRJNA400089 [60]), Lake Hazen, Nunavut (PRJNA430127 [61]), Jinchuan River (PRJNA437688, PRJNA437695, PRJNA437692, and PRJNA437697 [62]), Lake Neusiedl (PRJNA507590 [63]), Hydrocarbon Biodegradation (PRJNA523725 [64]), Bay of Biscay (PRJEB35647 [65]), Aldabra Atoll (PRJNA611521 [66]), Maozhou River Hyporheic Zone (PRJNA616197 [67]), Estuarine Sediment Mesocosms (PRJNA639965 [68]), Hydrodynamic Experiments (PRJNA777293 [69]), and Cardinal Pond, Poland (PRJNA832534 [70]) Soil Dabu Town, Hunan Province (PRJNA361046 [71]), Urban Wetland Soils (PRJNA415514 [72]), Tengger Desert Biocrusts (PRJNA543295, PRJNA647192, and PRJNA647699 [73]), Paddy Soil Water Management (PRJNA564714 [74]), Anaerobic Soil Disinfestation (PRJNA575041 [75]), East Asia Paddy Soil (PRJNA607877 [76]), Rice Rhizosphere (PRJNA611687 [77]), Georgia Salt Marshes (PRJNA666636 [78]), Paddy Soil Amendments 1 (PRJNA690162 [79, 80]), Maize Intercropping, Poland (PRJNA725644 [81]), Paddy Soil Amendments 2 (PRJDB12684 [82]), and Soil–Water Interface (PRJNA826420 [83]). Metaproteome comparisons Manus Basin Inactive Chimney (PRJEB27164 [84]), Manus Basin Active Chimneys (PRJEB5213 [85]), Soda Lake Biomats (PRJNA377096 [86]), Mock Communities (PRJEB19901 [86]), and Saanich Inlet (PRJNA247822 [87]). a The name of each data set is followed by the BioProject accession number(s) and one or two literature references; the second reference, if present, is for redox potential and/or O2 data. Bold text is used for laboratory or mesocosm experiments. Abbreviations: AMD, acid mine drainage; CROMO, Coast Range Ophiolite Microbial Observatory; nZVI, nanoscale zero-valent iron. O2 concentration and redox potential, the latter assumed to be stated relative to SHE, were taken from the primary publications (see Table S1 for details). Where needed, values were extracted from figures using g3data (https://github.com/pn2200/g3data, accessed on 11 October 2022). Temperature (T) and pH were also tabulated if available; otherwise their values were assumed to be 25°C and pH 7 for the conversion to Eh7 (see below). Where there are extreme values of temperature and pH, such as in hot springs and hyperalkaline systems, these variables are routinely reported; for the few data sets for non-extreme environments where neither T nor pH was reported (river & seawater: 1, lake & pond: 1, sediment: 4, soil: 2; see Table S1), their effects on Eh7 are likely to be relatively small. Brief descriptions of samples (e.g., substrate type, sample location, or collection time) were also recorded as applicable to each data set. Correction of Eh to pH 7 Because of the dependence of electrochemical potential on pH, known as the Nernst slope or Nernst factor, reported Eh values were corrected to pH 7 with (88) (1) Eh7=Eh+dEhdpH×(7−pH) where Eh7 denotes the corrected value and dEh/dpH is the theoretical slope, which is equal to −59.16 mV/pH unit at 25°C. Values of the Nernst slope at other temperatures were calculated from −2.303 RT/F, where R, T, and F are the gas constant, temperature in Kelvin, and Faraday constant. The theoretical relationship between pH and Eh as expressed in Equation 1 may not hold for all natural systems (89), so correcting Eh to pH 7 is not recommended for reporting primary results (90). Unit conversion of oxygen concentrations O2 concentrations reported in units of milligram per liter were converted to micromoles per liter using the molar mass of O2 (31.9988 g/mol). Concentrations reported as percent saturation (%) were converted to micromoles per liter by first combining Henry’s constant for O2 at the reported sample temperature, or 25°C if not available, with the partial pressure of O2 in the atmosphere at sea level (0.21 atm = 0.2128 bar) to calculate saturated O2 concentration, which was then multiplied by the percentage value to obtain concentration in micromoles per liter. Henry’s constants as a function of temperature were calculated using the subcrt function in the CHNOSZ R package version 1.4.3 (11). Sequence data processing One data set in our compilation is for environmental 16S amplicon rRNA (i.e., not shotgun) sequences reverse-transcribed to cDNA (Soil–Water Interface [83]); it was processed in the same way as the other data sets, but taxonomic abundances were derived from the taxonomic classifications of cDNA sequences. All other data sets are for 16S rRNA gene sequencing of environmental DNA; if results for cDNA were also reported, they were not included in our compilation. Three large data sets were subsampled so that they would not dominate the global analysis. The complete New Zealand Hot Springs data set consists of 925 samples (15); we used data for 42 randomly selected acidic samples (pH < 3) and 39 randomly selected circumneutral to alkaline samples (pH > 6). For the Bay of Biscay data set (65), we used data for the 47 sediment samples with redox potential measurements collected in 2017, which is a subset of the 256 samples collected in different years. For the Port Microbes data set (14), we used data for the first three prefilter and first three postfilter samples with at least 20,000 sequenced read pairs and complete metadata for each port (114 of 1,373 available samples). Sequence data were processed with a custom pipeline consisting of merging of forward and reverse reads where applicable, quality and length filtering, singleton removal, reference-based chimera detection, and taxonomic classification. The R script implementing the pipeline is available (see Data availability). The pipeline uses the following software and databases: fastq-dump from the NCBI SRA Toolkit for generating FASTQ files from downloaded SRA files, VSEARCH (91) for merging, singleton and chimera detection, seqtk (https://github.com/lh3/seqtk, accessed on 26 April 2023) for extracting non-singletons, SILVA SSURef NR99 database version 138.1 (92) for chimera detection, and RDP Classifier version 2.13 (93) for taxonomic classification with the included training set (RDP 16S rRNA training set no. 18 07/2020) and the default confidence threshold of 80%. For the two 454 pyrosequencing data sets in our compilation (27, 45), length filtering was specified as a minimum read length of 300 and maximum length of 500 (-fastq_minlen and -fastq_maxlen options of VSEARCH); this setting is based on the processing steps described by Kocur et al. (45). To process sequence data from the Illumina and Ion Torrent platforms, we adjusted the truncation length (-fastq_trunclen option of VSEARCH) to maximize the performance of the taxonomic classifier. Specifically, a truncation length was individually selected for each data set as the highest number, rounded to the nearest 10, where the majority of reads passed the filtering step. After length and quality filtering and singleton removal, runs were subsampled to a depth of 10,000 sequences before chimera detection and taxonomic classification; the subsampling was done only to reduce processing time and is not the same as normalization of sequencing depth used for computing diversity metrics (see Methodological limitations and justification). Settings for each data set, sequence processing statistics, and additional details are listed in Table S2. Taxonomic mapping and amino acid composition of community reference proteomes RDP classifications at the root or domain level or to chloroplast or eukaryota were omitted. No filtering of mitochondrial reads was done because no sequences labeled as such are present in the training set included with the RDP Classifier. The lowest assigned taxonomic level (from genus to phylum) for each remaining sequence was mapped to the NCBI taxonomy represented in RefSeq by automatic matching of both rank and taxon names or by manual mapping for particular taxa as described previously (5). The following within-level manual mappings were used (RDP → NCBI): genus, Escherichia/Shigella → Escherichia, Gp1 → Acidobacterium, Gp6 → Luteitalea, GpI → Nostoc, GpIIa → Synechococcus, GpVI → Pseudanabaena; family, Family II → Synechococcaceae, Ruminococcaceae → Oscillospiraceae; order, Clostridiales → Eubacteriales, Rhizobiales → Hyphomicrobiales; class, Planctomycetacia → Planctomycetia; phylum, Cyanobacteria/Chloroplast → Cyanobacteria. Several cross-level manual mappings were also used: genus Subdivision3_genera_incertae_sedis → family Verrucomicrobia subdivision 3, genus Spartobacteria_genera_incertae_sedis → class Spartobacteria, class Cyanobacteria → phylum Cyanobacteria, and class Actinobacteria → phylum Actinobacteria. No manual mapping was done for archaeal taxa. Percentages of mapped taxa are listed in Table S2. Unmapped taxonomic assignments were omitted from subsequent analysis. For sequencing runs for libraries generated using domain-specific bacterial or archaeal primers, only the taxonomic assignments in the respective domain were kept. For studies that used universal primers, the RDP Classifier assignments within each domain were selected. Samples with less than 100 sequences with lowest-level taxonomic assignments from genus to phylum level for either bacteria or archaea were excluded from downstream analysis. If the archaeal metacommunity in any data set was represented by fewer than four remaining samples, it was not processed further. We used previously compiled reference proteomes (5) for archaeal and bacterial taxa at levels from genus to phylum derived from the NCBI RefSeq database release 206 (94). To generate the reference proteomes, only species-level NCBI taxonomic IDs (taxids) were used, and archaeal and bacterial species with less than 500 reference protein sequences were excluded. For each species-level taxid, the sum of amino acid compositions of all reference sequences was divided by the number of reference sequences to obtain the mean amino acid composition; this was done so that species with different proteome sizes contribute equally to the reference proteomes of higher-level taxa. For each genus, the mean amino acid compositions of all species-level taxa in that genus were summed and divided by the number of species to obtain the amino acid composition of the reference proteome. Analogously, the mean amino acid compositions of all species in each family, order, class, and phylum were averaged to obtain the reference proteomes for taxa at those levels. Within each domain, the counts of all mapped lowest-level taxonomic assignments were multiplied by the amino acid compositions of the corresponding reference proteomes and divided by the total count of mapped assignments to obtain the amino acid composition of the community reference proteome. Metaproteomes Publicly available processed metaproteomic data were used. Except as noted below, protein IDs in pep.xml or mztab files, excluding sequences labeled as contaminants and reverse decoy sequences, were matched to the protein reference database for that study to get protein sequences. For pep.xml files, the first Protein ID for each mass spectrum search query was used; protein abundance was estimated by counting multiple identifications of the same protein in different peptide spectra. For mztab files, the amino acid composition of each identified protein was multiplied by the number of peptide spectral matches. The amino acid compositions of all proteins were then summed and used to calculate Z C. The ProteomeXchange accession numbers and specific files used are listed next. Manus Basin Inactive Chimney (84): PXD010074; protein sequence database: DEAD_Chimneys_20170720_nr_fw_rev_cont.fasta; protein IDs: Mudpit_121204_V2_P6_CH_SM_Chimney_3b_1.pride.mztab.gz. Manus Basin Active Chimneys (95): PXD009105; protein sequence database: 140826_StM_Chimney39surface53surface_fw_rev_cont.fasta; protein IDs: Mudpit_121026_V2_P6_AO_SM_Chimney_1a_10.pride.mztab.gz (diffuse-flow chimney RMR-D) and Mudpit_121028_V2_P6_AO_SM_Chimney_2a_1.pride.mztab.gz (focused-flow chimney RMR5). Soda Lake Biomats (86): PXD006343; protein sequence database: SodaLakes_AllCombined_Cluster95ID_V2.fasta; protein IDs: GEM-(01).pep.xml and LCM-(01).pep.xml. Mock Communities (86): PXD006118; protein sequence database: Mock_Comm_RefDB_V3.fasta; protein IDs: 12 pep.xml files (three mock communities with four replicates each for 260-min 1D-LC-MS/MS runs). Saanich Inlet (87): PXD004433; protein sequence database: SaanichInlet_LP_ORFs_2015-01-13_Filtered.fasta; protein IDs: SBI_Metagenome2015_AllProteinsAllExperiments.txt (the ScanCount column in this table was used to estimate protein abundance). Metaproteomic experiment names were mapped to sample names using Table S2 of reference 87. To obtain community reference proteomes for comparison with metaproteomes, 16S rRNA gene sequencing data sets were taken from the studies cited above, except for Manus Basin Active Chimneys (85). Computational methods Average oxidation state of carbon (Z C) for proteins can be calculated from (12) (2) ZC=−h+3n+2o+2sc where the lower case letters c, h, n, o, and s are the numbers of the respective elements in the elemental formula. Rather than using Equation 2 directly, Z C was calculated by combining amino acid compositions of community reference proteomes with precomputed Z C values for amino acids; the calculation includes weighting by the number of carbon atoms in each amino acid (12). Carbon oxidation state was computed from amino acid composition using the ZCAA function in the canprot R package version 1.1.2 (https://cran.r-project.org/package=canprot), and linear regressions were performed using lm in R version 4.2.1 (96). The world map, drawn with the Winkel Tripel projection, was made using the oce R package version 1.7-2 (97) and its included coastlineWorld data set together with shapefiles for the North American Great Lakes (98). Statistics We subjected one hypothesis to multiple trials by assessing correlations between Eh7 and Z C obtained from different local and global data sets. Alternative hypotheses using Eh or O2 concentration instead of Eh7 were only tested for the pan-environmental global data set. Except for the subsampling of samples and sequencing reads described above, no data were removed from the analysis; in particular, outliers were not removed. The statistical analyses used are the slopes of linear regressions (m) and Pearson correlation coefficients (r). Values of the slope are reported with margin of error (MOE) as m ± MOE 95, where [m − MOE 95, m + MOE 95] is the 95% CI; this was used in lieu of P-values for correlations. To assess the statistical significance of the frequency of positive results, P-values were calculated using exact one-sided binomial tests assuming a 50% chance of a positive correlation. Methodological limitations and justification The assignment of entire data sets to environment types may include samples that do not all fit that description. For example, some sediment data sets also contain samples of overlying water (see Table S1). A more fine-grained approach would be to describe environment types at the level of samples rather than data sets. However, for the purposes of local-scale comparisons, including all samples in each data set is more faithful to the source study designs, some of which include different sample types that are used as controls. We feel this advantage outweighs the negative impact on the global-scale comparison, especially since the number of potentially misclassified samples in the global analysis is relatively small. We did not use operational taxonomic units (OTUs) or amplicon sequence variants (ASVs) in our analysis. A common reason for using OTUs or ASVs is to generate diversity metrics, including various metrics for alpha and beta diversity. In contrast, the calculation of chemical metrics only requires estimates of total elemental composition at the community level. A straightforward approach is to add up the elemental compositions of the reference proteomes corresponding to each taxonomically classified sequence without clustering. Because the most abundant taxa dominate the elemental composition, variability of sequencing depth among data sets is not a major source of uncertainty in the calculation of chemical metrics. We used the default RDP Classifier training set for taxonomic classification rather than the more popular SILVA database. The SILVA database is larger and might provide greater taxonomic resolution (i.e., a higher proportion of genus-level assignments). Lower taxonomic resolution tends to lessen the differences of Z C (see Fig. S2 in reference 5); therefore, this limitation is more likely to generate false-negative than false-positive results. There are limitations inherent in the use of reference proteomes and taxonomic mapping; reference proteomes contain no information about protein expression levels, and mapping between the RDP and NCBI taxonomies is somewhat uncertain. To overcome the limitations of a taxonomy-based approach, future work should consider whether phylogenetic inference methods, which are currently available for inferring functional traits (99), could be modified to infer amino acid composition. However, phylogenetic inference requires a reference database of traits, so limited knowledge of protein expression levels would still be an issue. RESULTS Theoretical background: incomplete metastable equilibrium The theoretical relation between environmental variables and molecular oxidation state can be quantified by using methods adapted from geochemical thermodynamics. Under ambient conditions, proteins are thermodynamically unstable relative to smaller molecules and tend to spontaneously (i.e., exergonically) hydrolyze into amino acid monomers, which themselves are unstable relative to inorganic compounds. By excluding from consideration the potential for the formation of amino acids and other small molecules, a metastable equilibrium model for reference proteomes can be constructed. The model described here is defined in terms of thermodynamic components (see reference [100]), which are mathematical abstractions that theoretically relate the elemental composition of reference proteomes to chemical variables such as Eh and pH but do not necessarily reflect actual mechanisms of protein synthesis or evolution. For the purposes of illustration, we selected reference proteomes for five methanogen species (Methanococcus vannielii, Methanococcus maripaludis, Methanococcus voltae, Methanobrevibacter smithii, and Methanofollis liminatans) having a range of Z C from −0.220 for M. vannielii to −0.154 for M. liminatans (see Fig. 1 in reference 101). These species represent a diverse phylogeny, including both Class I and Class II methanogens, but they are all mesophiles with optimal growth temperatures of 30°C–40°C. We wrote overall formation reactions for the reference proteomes from thermodynamic components, calculated the Gibbs energies of the reactions as a function of pH and Eh at 25°C, then identified the reference proteome with the lowest Gibbs energy to plot relative stability fields (Fig. 1a; for further details, see Materials and Methods: Thermodynamic calculations). Because the Gibbs energies are positive except at the most reducing conditions (near the stability limit of water), the reference proteomes are unstable relative to the thermodynamic components across most of the diagram, and “relative stability” denotes the least unstable proteome compared to the others. Fig 1 Thermodynamic model for the relationship between carbon oxidation state of reference proteomes and redox potential. (a) Eh–pH diagram for reference proteomes of five selected methanogen species. All proteomes are unstable relative to the basis species used in the calculations (glutamine, glutamic acid, cysteine, H2O, H+, and e -); the relative stability fields, therefore, represent the proteome that is least unstable at particular values of Eh and pH. The gray area represents reducing conditions that are beyond the stability limit of water. (b) Carbon oxidation state of relatively stable reference proteomes as a function of Eh or Eh7 at three pH values. The upper plot shows Z C of the most stable reference proteome at three pH values, corresponding to the vertical dashed lines in (a). Note, for instance, that the reference proteome for Mvo, which is relatively stable at pH = 5, has Z C that is intermediate between those of Mma and Msm, which are relatively stable at pH = 9. The lower plot shows the same profiles after correction to Eh7 using equation 1. (c) A simple conceptual model for equilibrium and non-equilibrium effects on the relation between carbon oxidation state and Eh7. The first plot shows theoretical equilibrium values of Z C sampled at equal intervals of Eh7, taken from the profile for pH = 7 in (b). The second plot shows Z C for randomly selected archaeal and bacterial species from RefSeq, representing non-equilibrium biological variability; this particular choice of random species displays essentially zero correlation in the plot. The third plot shows a linear combination of the first and second plots. Abbreviations: Mli, Methanofollis liminatans; Mma, Methanococcus maripaludis; Msm, Methanobrevibacter smithii; Mva, Methanococcus vannielii; Mvo, Methanococcus voltae. The reference proteomes with the lowest and highest Z C are relatively stable at low and high Eh, respectively, and cross sections of the Eh–pH diagram at discrete values of pH exhibit a stepwise increase of Z C with increasing Eh (Fig. 1b). Because of the negative slopes of the stability-field boundaries on the Eh–pH diagram, lower pH is associated with higher Eh, and vice versa. To allow comparing observations at different pHs, a quantity known as Eh7 can be used (88); it is calculated by moving along a line of constant slope from a given pH to pH = 7 (equation 1). Values of Z C in the metastable equilibrium calculation are widely dispersed along the Eh scale at three pHs from 5 to 9 but overlap when plotted against Eh7 (Fig. 1b). It should be noted that the relative stabilities of proteomes depend on both pH and Eh; the correction to Eh7 is independent from the metastable equilibrium calculation and is only used to illustrate how values of Eh obtained at different pH values can be compared. Real biological systems are not at complete metastable equilibrium. Chemical adaptation is simply the hypothesis of an incomplete equilibrium. We simulated non-equilibrium biological effects by assigning Eh7 values at equal intervals, then picking a random sample of species from the RefSeq database for which the slope of the regression between the assigned values of Eh7 and Z C of the species’ reference proteomes is close to zero (Fig. 1c). Then, we calculated a weighted mean of Z C assuming weights of 20% for the metastable equilibrium values and 80% for those of the randomly sampled species. The addition of a random effect decreases the slope and widens the confidence interval of the slope. The slope calculated in this simulation (0.133/V) is comparable to the highest slopes empirically observed for local-scale data sets below. This simple simulation highlights the feasibility of detecting incomplete metastable equilibrium against a background of random biological variation. To generalize from this model, thermodynamic considerations predict a positive association between Z C and Eh7. The aim of this study is to empirically test this prediction using data for natural communities, which can unveil not only positive but also negative or non-significant correlations. The analysis below uses data for modern communities and genomes, but our interpretation is about evolutionary differences, similar to other comparative genomic studies (102). Although the hypothesis of chemical adaptation presupposes a sufficient period of genome evolution under steady redox conditions, the prediction may also be applicable to short-term experimental interventions because of the phenomenon of species sorting—that is, environmental selection for preadapted species (103). Z C of reference proteomes reflects oxygen tolerance and is distinct from metaproteomes Various microorganisms have distinct oxygen tolerance. The “List of Prokaryotes According to Their Aerotolerant or Obligate Anaerobic Metabolism” (version 1.2) (104) classifies 644 prokaryotic genera as strictly anaerobic or aerotolerant, including facultative anaerobes. This list is predominated by bacteria, with two archaeal genera present in RefSeq (Methanobrevibacter and Methanosphaera). We found that the carbon oxidation state of reference proteomes for aerotolerant genera is significantly higher compared to strict anaerobes (Fig. 2a). This finding is broadly consistent with the notion that genomes code for protein sequences that are to some extent chemically adapted to redox conditions. Fig 2 Z C of reference proteomes compared with oxygen tolerance and with metaproteomes. (a) Carbon oxidation state (Z C) of reference proteomes for strictly anaerobic and aerotolerant genera in the RefSeq database. The oxygen tolerance of prokaryotic genera was taken from Table S1 of reference (104). Of the genus names in that table, 64 were not matched to RefSeq and were omitted from the comparison. Student’s two-sided t-test was used to calculate P value. (b) Comparison of metaproteomes and community reference proteomes. Z C was calculated for all proteins identified in each metaproteome. Independently, Z C was also calculated for community reference proteomes derived from 16S rRNA sequences and RefSeq proteomes. The dashed line is the theoretical 1:1 line; the solid line shows linear regression of all data points with Pearson correlation coefficient (r) and slope (m) indicated in the legend. Next we aimed to find out whether estimates of carbon oxidation state for community reference proteomes are similar to values calculated for actual metaproteomes. The data sets were selected with the criteria of joint availability of metaproteomic data (including a protein sequence database) and 16S rRNA gene sequences. The analyzed data sets include active and inactive chimneys at the Manus Basin hydrothermal vent field (84, 85), the water column of the Saanich Inlet in British Columbia, Canada (87), and soda lake biomats from the Rocky Mountains in Canada (87). We also analyzed data for mock communities generated in the latter study (87); although the species composition of the mock communities was known, we did not use that information but instead generated the community reference proteomes from 16S rRNA gene sequences. The proteins identified in the metaproteomic studies were used to calculate Z C (see Materials and Methods: Metaproteomes). In parallel, 16S rRNA gene sequences for the same samples were used to generate community reference proteomes. Because of limitations in mapping Ribosomal Database Project (RDP) classifications to the NCBI taxonomy (see Materials and Methods: Taxonomic mapping and amino acid composition of community reference proteomes), a reference proteome for each identified taxon in a community was not always available. The percentages of mapped taxa are listed in Table S2. The community reference proteomes were calculated by multiplying the amino acid composition of the reference proteome for each identified taxon by the taxonomic abundance and omitting those taxa without available reference proteomes. There is a weak positive correlation between metaproteomes and community reference proteomes, but a generally higher Z C of metaproteomes (Fig. 2b). In experiments designed to provide an unbiased estimate of membrane protein abundance, the total abundance of cytoplasmic proteins in Escherichia coli is more than double than that of membrane proteins (data from Table S3A of reference 105). Furthermore, common mass spectrometry (MS)–based proteomics workflows have limited ability to detect membrane proteins (106). Because they are enriched in hydrophobic amino acids, membrane proteins tend to be more reduced than cytoplasmic proteins (107). For instance, in the yeast Saccharomyces cerevisiae, sequences of proteins localized to the cytoplasm and plasma membrane have mean Z C of −0.127 and −0.188, respectively (107). Therefore, technical bias against detecting membrane proteins in MS-based proteomics and the higher natural abundance of cytoplasmic than membrane proteins are two factors that could possibly contribute to the higher Z C of proteins observed in metaproteomes compared to community reference proteomes. An interesting result is that active hydrothermal chimneys have the most reduced community reference proteomes among the data sets compared in Fig. 2b. This result shows that the highly reducing environment associated with hydrothermal fluids corresponds to relatively reduced community reference proteomes. Notably, metaproteomes indicate that the expressed proteins of communities in active chimneys are also reduced relative to those in most other environments. Unfortunately, metaproteomic data were not available for assessing correlations with redox potential in the remainder of this study. Winogradsky columns Based on the above considerations, we predicted to find a positive correlation between redox potential and carbon oxidation state of community reference proteomes for various environments (Fig. 3a). We first analyzed data for Winogradsky columns, which are a type of microcosm constructed by placing water and sediment, enriched with a carbon source, into a sealed transparent container and allowing it to develop under controlled lighting and temperature conditions (16). As far as we know, no published studies have reported 16S rRNA gene sequences and redox potential measurements along depth profiles in the same Winogradsky columns, so we used data from different sources to represent general features. Fig 3 Overview of methods and chemical depth profiles in different Winogradsky columns. (a) Schematic overview of data and methods used in this study. (b) Measurements of oxidation–reduction potential (denoted here as Eh; dashed line) in a Winogradsky column made with acidic sediment, taken from Fig. S5 of Diez-Ercilla et al. (17) with the depth scale adjusted so the sediment–water interface is at 0 cm. Values of Eh7 computed from Equation 1 are also shown (solid line). (c) Values of Z C for community reference proteomes computed in this study using 16S rRNA gene sequences reported by Rundell et al. (16) for Winogradsky columns composed of non-acidic sediment from ponds in Massachusetts, USA (BioProject PRJNA234104). The box-and-whisker plot represents data for samples collected from the same depth intervals in different columns. At each depth, N is the number of samples, and the center line, box width, whiskers, and points denote the median, interquartile range (IQR), the most extreme values within 1.5× IQR, and values outside this range. “Top” indicates samples collected by scraping the biofilm on the top surface of the sediment and “SWI” indicates samples of the sediment–water interface collected by drilling into the side of the column (16). Figure 3b shows Eh measurements taken from reference (17) for a column consisting of acidic sediment and water from pit lakes in the Iberian Pyrite Belt, Spain. Because of the acidic conditions in these experiments, values of Eh7 calculated with Equation 1 are shifted to lower values than the reported redox potential, but they still exhibit a decreasing trend with depth. Redox potential also decreases with depth in the sediment of mature Winogradsky columns prepared using sediment from non-acidic ponds (108, 109), and sediments have been reported to be more reducing than the overlying water (see Fig. 5 in reference [110] and Fig. S1A in reference 111). It follows that a decrease in redox potential with depth is a general feature of mature Winogradsky columns. In Winogradsky columns made with non-acidic sediment from ponds in Massachusetts, USA, relatively high abundances of Alphaproteobacteria and Betaproteobacteria occur near the tops of the columns, and more abundant Firmicutes, particularly Clostridia, are found in the lower parts (16). Proteobacteria and Clostridia are lineages with relatively oxidized and reduced reference proteomes, respectively (5), and changes in the relative abundances of members of these taxa are the primary drivers for the decreasing trend of Z C with depth shown in Fig. 3c. In parallel with the relatively low similarity of communities in replicate samples for the SWI as revealed by UniFrac distances (16), the SWI has the greatest variability of Z C, which further illustrates that Z C is a chemical representation of the taxonomic composition of communities. Although the plots in Fig. 3b and c may be visually compelling, scatterplots are needed to make direct comparisons of Z C and Eh7 rather than comparing each with a third variable (e.g., depth). We did not make a scatterplot for Winogradsky columns because of the lack of paired redox potential and 16S rRNA data for the same columns. Local-scale analysis We analyzed 68 data sets compiled from 66 studies that reported both environmental 16S rRNA gene sequences and redox potential measurements for samples from environments categorized here as river & seawater, lake & pond, geothermal areas, hyperalkaline fluids, groundwater, sediment, and soil. The data sets in our compilation have a worldwide distribution (Fig. 4a). Compared to the ranges of natural environments described by Baas Becking et al. (90), our analysis includes samples with higher pH in hyperalkaline fluids, lower Eh at near-neutral pH in some soils and freshwater systems, and lower Eh at acidic conditions in some geothermal areas (Fig. 4b). For each data set, Eh7 and Z C of the bacterial community reference proteome (referred to hereafter as bacterial Z C) were used to make a scatterplot (Fig. S1). The plot legends include the number of samples (N), Pearson correlation coefficient (r), and slope (m). The units of slope (1/V) come from the dimensionless y variable (Z C has no units because it is a ratio of elemental abundances) divided by the x variable with units of V. If sufficient numbers of taxonomically classified and mappable archaeal sequences were available (see Materials and Methods), a second plot for archaeal Z C was also made. The selected plots in Fig. 5a show that geographically separated sediment samples from the Bay of Biscay, Spain (65), and sediment and soil samples in Hunan Province, China (71), exhibit positive correlations between Eh7 and bacterial Z C, but samples from different depths of sediment cores in Daya Bay, China (60), have a negative correlation. Fig 4 Sample locations and Eh–pH diagram. (a) Sample locations were obtained from NCBI BioSample metadata or primary publications. If the location of source material for laboratory or mesocosm studies was not specified, the institutional address was used. Small open circles indicate sampling transects for paddy soils in East Asia (76), hot springs in the Southern Tibetan Plateau (38), and water samples from the Three Gorges Reservoir (25); small filled triangles represent 20 worldwide ports (14). (b) Eh–pH diagram for all environment types. The outline is redrawn from Baas Becking et al. (90) and represents the range of natural environments described in that study. Note that this figure shows values of Eh, not Eh7 as in the following figures. Fig 5 Associations between Eh7 and Z C at local scales. (a) Selected data sets for sediments and soil. Scatterplots between Eh7 and bacterial Z C are shown with linear regressions. Legends indicate the number of samples (N), Pearson correlation coefficient (r), and slope of the linear regression (m) ± the margin of error for the 95% confidence interval. (b) Slopes of linear regressions plotted against the decimal logarithm of numbers of samples in each data set for bacterial communities. The data sets with the largest effect sizes (i.e., those for which the absolute value of slope is > 0.1/V) are indicated by larger points. Of these nine data sets, eight have positive slopes; the only one with a negative slope is the data set for Daya Bay shown in (a). (c) Linear regressions for bacterial and archaeal communities in geothermal areas. Colors are used to represent sample characteristics (acidic water, circumneutral to alkaline water, and sediment). Numbers for data sets are (1) acidic and (2) circumneutral to alkaline New Zealand hot springs (15), (3) Eastern Tibetan Plateau (36), (4) Uzon Caldera (including nine samples for acidic water and sediment and one high-pH sample) (37), and (5) Southern Tibetan Plateau (38). The correlations between Eh7 and bacterial Z C are positive for each of the five data sets for geothermal areas and have some of the highest values of slope for all environments (Fig. 5b), but the ranges of Z C differ widely depending on fluid and sample characteristics (Fig. 5c). Circumneutral to alkaline water samples from hot springs in the Eastern Tibetan Plateau (36) and New Zealand (15) have the lowest ranges of bacterial Z C not only for hot springs but also for all data sets in our compilation. Archaeal Z C exhibits both positive and negative correlations with Eh7 in geothermal areas (Fig. 5c), so the hypothesized thermodynamic effect on chemical differences appears to be stronger for the bacterial domain. After geothermal areas, soils have the next highest frequency of positive correlations (10 out of 12 data sets or 83%). We analyzed seven soil data sets from laboratory or mesocosm studies in which treatments included water management (flooding and draining), amendment with organic or silicon-rich compounds, and anaerobic soil disinfestation; all of these data sets yield a positive correlation between Eh7 and bacterial Z C (see Table S1). The analysis also includes a data set for millimeter-scale variations of soil communities below the soil–water interface (83); because the sequences in this data set were obtained from environmental 16S rRNA rather than 16S rDNA, the carbon oxidation state of reference proteomes of active and not only present community members is evidently aligned with the redox gradient. Global-scale analysis A global analysis performed by combining individual data sets yields a positive slope for bacterial communities in each of the seven environment types (Fig. 6a). The data sets for groundwater show the strongest global correlation (r = 0.43), followed by lake & pond and soil (both with r = 0.37). Although geothermal areas exhibit strong positive correlations for bacterial communities at local scales (Fig. 5c), the differences between alkaline water and acidic water and sediments blur the pattern at a global scale. Owing at least in part to the aggregation of local-scale data sets to make the global compilation, there is a relatively large amount of scatter in the global comparison for each environment type, and it would be unwise to try to predict Z C for any particular community from the global correlation with redox potential. In contrast to bacteria, archaea show a negative global correlation in each environment type. Combining data for all environments yields positive and slightly negative correlations between Eh7 and Z C for bacteria and archaea, respectively, but only the correlation for bacteria is statistically significant (i.e., the 95% confidence interval of the slope for bacteria is entirely greater than zero, while that for archaea crosses zero; Fig. 6b). Fig 6 Global-scale associations between Eh7 and Z C. (a) Sample values and linear regressions for bacterial and archaeal communities in each environment type. The legend text is grayed out for environments represented by fewer than five data sets with archaeal sequences analyzed in this study. (b) Pan-environmental comparison. The legends indicate the number of data sets (top left), number of samples and slope of the linear regression ± the margin of error for the 95% confidence interval (top right), and Pearson correlation coefficient (bottom right). Our compilation of metadata includes measurements of oxygen concentration, if available, for each data set. O2 concentrations reported as zero or below the detection limit were set to zero; samples without reported O2 measurements were not included in the analysis described next. Although a positive correlation between bacterial Z C and O2 is evident in Fig. S2 (N = 1,493, r = 0.17), it is weaker than the correlation between Z C and Eh7 for the same set of samples (r = 0.37). Furthermore, bacterial Z C is more strongly correlated with Eh7 than with values of Eh that have not been corrected to pH 7 (r = 0.29). For the 256 samples in our compilation with both O2 measurements and archaeal 16S rRNA gene sequences, the correlation with Z C is significant and positive for both Eh7 (r = 0.17) and O2 (r = 0.44). This suggests that geochemical influences on archaeal protein evolution may be more strongly mediated by oxygen (explaining, in part, the lack of correlation for archaea in Fig. 6b), in contrast to an apparently stronger influence on bacterial proteins by redox potential. Analysis of data for one primer set All primer sets for the 16S rRNA gene exhibit amplification bias, which refers to the variable efficiency of amplification of DNA from different taxa (112), and our global-scale analysis could in principle be affected by primer set–specific differences of amplification bias. To control for differences of amplification bias among primer sets, we performed a second global analysis including only data sets that were generated using the 515F/806R primer set adopted for use by the Earth Microbiome Project (113) (see Table S1). This analysis reveals positive correlations between Eh7 and bacterial Z C in all environment types except river & seawater and lake & pond (Fig. S3a). Furthermore, the correlation for bacteria in the pan-environmental comparison is stronger when considering only the 515F/806R primer set (N = 1074, r = 0.32, slope = 0.023 ± 0.004/V; Fig. S3b) than mixed primer sets (N = 2792, r = 0.28, slope = 0.016 ± 0.002/V; Fig. 6b). Therefore, limiting the analysis to data sets generated using one primer set supports our main conclusions. Statistical analysis Analyses of multiple data sets represent independent tests of a single hypothesis, so we used binomial probabilities to simulate the number of successes from a certain number of trials. The frequency of observed positive correlations at the local scale for bacterial communities in soil and geothermal environments could occur by chance with probabilities (binomial P-values) of 0.019 and 0.031, respectively, and the combined outcome for all environments has binomial P = 0.00007 (Table 2a). If the local-scale results for bacterial communities are filtered to include only slopes that are significantly negative or positive (i.e., those for which the minimum and maximum values in the 95% confidence interval have the same sign), then only five data sets have significantly negative slopes (two for lake & pond and three for sediment environments), and the occurrence of 31 significantly positive correlations from 36 data sets for all environments has binomial P = 0.00001 (Table 2b). TABLE 2 Tally of regression slopes for local-scale correlations and probabilities that the results could occur by chance (P-values) a a Bacteria Archaea N tot N pos N neg P N tot N pos N neg P River & seawater 10 6 4 0.377 4 4 0 0.062 Lake & pond 9 6 3 0.254 3 2 1 0.5 Geothermal 5 5 0 0.031 4 1 3 0.938 Hyperalkaline 6 5 1 0.109 3 2 1 0.5 Groundwater 10 7 3 0.172 6 2 4 0.891 Sediment 16 11 5 0.105 6 2 4 0.891 Soil 12 10 2 0.019 5 2 3 0.812 Total 68 50 18 0.00007 31 15 16 0.640 b Bacteria Archaea N tot N pos N neg P N tot N pos N neg P River & seawater 3 3 0 0.125 2 2 0 0.25 Lake & pond 5 3 2 0.5 1 1 0 0.5 Geothermal 3 3 0 0.125 0 0 0 NA Hyperalkaline 3 3 0 0.125 3 2 1 0.5 Groundwater 3 3 0 0.125 2 0 2 1 Sediment 10 7 3 0.172 3 0 3 1 Soil 9 9 0 0.002 3 1 2 0.875 Total 36 31 5 0.00001 14 6 8 0.788 a (a) All data sets; (b) data sets filtered to include only those with statistically significant slopes. P is the probability for N pos or more successes to occur by chance in N tot trials with a 0.5 hypothesized probability of success in each trial. Nneg, number of data sets with negative regression slope; Npos, of data sets with positive regression slope; Ntot, total number of data sets analyzed;. A 95% confidence interval that is entirely greater than zero indicates that the slopes in the global analysis are significantly positive for each environment type except for river & seawater (Fig. 6a). The correlation for river & seawater can be considered borderline significant because the 95% confidence interval of the slope extends down to but does not cross zero. Collectively, the occurrence of six or more significantly positive correlations for bacterial communities in seven environment types has binomial P = 0.0625. Furthermore, of the nine local-scale data sets for which the absolute value of the slope is greater than 0.1/V, eight have positive slopes (Fig. 5b). In other words, a predominance of positive slopes is exhibited by the subset of data sets with the largest effect size and not only by all data sets in the compilation for bacterial communities. DISCUSSION This study is primarily a contribution toward geochemical biology, which is a research area concerned with the effects of geochemistry on genome evolution. To test the hypothesis that protein sequences coded by genomes in microbial communities are chemically adapted to redox conditions, we assessed the sign of correlations between Z C of community reference proteomes and environmental redox potential corrected to pH 7 for more than 60 publicly available data sets. We observed positive correlations for a majority of data sets for bacterial communities at local scales (Table 2) and at a global scale for each of seven environment types (Fig. 6). Despite this, the regressed slopes vary widely (Fig. 5b), and there is a substantial spread around the regression line in most scatterplots. Therefore, redox conditions appear to have a modulating rather than controlling effect on the carbon oxidation state of bacterial protein sequences. A limitation of this study is that reference proteomes carry no information about protein abundance and cannot be used to make inferences about the effects of geochemistry on protein expression levels or whole-cell elemental composition. As proteins typically make up approximately half of the cellular dry weight, the elemental composition of cells is strongly related to protein expression. On the one hand, highly expressed proteins require fewer total mutations than lowly expressed proteins for stoichiometric constraints to be accommodated by sequence evolution. On the other hand, lowly expressed proteins evolve faster than highly expressed proteins (114). A tradeoff between abundance and evolvability suggests a reason for the sequences of both highly and lowly expressed proteins to be chemically adapted to environmental conditions and may explain in part why community reference proteomes exhibit correlations with redox conditions despite having assumed equal contributions from proteins that actually vary widely in abundance. Biological and technical biases represent additional limitations that demand a conditional rather than definitive conclusion. Horizontal gene transfer and large mobile genetic elements (e.g., plasmids) are not adequately represented by reference proteomes, and there are multiple sources of bias that affect the analysis of gene sequence data sets–from DNA extraction to sequencing to bioinformatic processing. These biases affect comparisons both within and between data sets; they do not cancel out even for the same sequencing method. Therefore, the findings of this study–and many others–are potentially invalid (115). Bias, instead of true biological variation, could explain why correlations for some data sets are stronger than others. Confirming these trends by using another method (e.g., shotgun metagenomes) would help give more confidence to these results, although bias is also problematic for shotgun metagenomic sequencing (115) and for metaproteomics analysis (116). Among data sets for extreme environments, hypersaline lakes in the Monegros Desert of Spain (34) are characterized by very high Z C values of the archaeal communities (see Fig. S1 and the plot for lake & pond data sets in Fig. 6a). This trend is consistent with highly oxidized proteins inferred from metagenomes in other hypersaline samples (12). In contrast, moderately alkaline hot springs have the lowest range of Z C for all environments considered here. The consistently positive correlations for geothermal areas despite differences in pH and sample type (Fig. 5c) suggest that bacterial metacommunities can attain a low-energy state with respect to redox gradients in a complex physicochemical milieu. However, redox potential does not predict the large offsets of Z C among circumneutral to alkaline water, acidic water, and geothermally heated sediment samples, so other chemical and biological factors probably drive the divergent evolution that is apparent in different types of hot spring samples. There is a striking contrast between negative correlations observed for some sediment cores (see Daya Bay in Fig. 5a and Mai Po Wetland and Lake Neusiedl in Fig. S1) and the mostly positive correlations for soil data sets. Soils exhibit low phylogenetic diversity despite high species-level diversity; in contrast, sediments have higher phylogenetic diversity than many other environments (117). These differences might imply a greater extent of redundancy in the elemental composition of reference proteomes for distinct species in soils, which could contribute to the higher frequency of positive correlations between Eh7 and bacterial Z C for soils compared to sediments. At a global scale, it is not sediments but river & seawater communities that exhibit the weakest correlation compared to other environments. In contrast to lakes, many of which exhibit vertical redox stratification, the distribution of Eh7 for river & seawater data sets is weighted toward mostly positive values (Fig. 6a), suggesting that the magnitude of the sampled redox gradients is less in those systems; this may reduce the power to detect an association, if there is one, between Eh7 and Z C. In the global pan-environmental comparison, the carbon oxidation state of reference proteomes for bacterial communities is more strongly associated with Eh7 than either Eh or O2 concentration. A significant association with Eh7 is absent from the global pan-environmental comparison for archaea (Fig. 6b), but limiting the analysis to samples with reported oxygen concentrations reveals a positive association for archaea that is stronger for O2 than for Eh7, in contrast to bacteria (Fig. S2). This could suggest that protein evolution in archaea is more sensitive to oxygen levels, whereas bacterial protein sequences may be stronger indicators of redox potential than oxygen. However, we consider the results for archaea to be more uncertain because of limitations of archaea-specific primers (e.g., see reference 32) and the absence from the RefSeq database of representatives of the DPANN superphylum such as Woesearchaeota and Pacearchaeota that are abundant in some environments (118). Across all data sets, the average percentages of genus-level classifications made by the RDP Classifier that were mapped to the NCBI taxonomy are 86% for bacteria and 77% for archaea, suggesting that archaeal communities are not as well represented by available reference proteomes. In addition, reference proteomes for archaea are likely to be lower quality because automated annotation pipelines may not be current with recently discovered translation mechanisms (119). Conclusions In the context of millions of years of sequence evolution, a geochemical thermodynamic model makes a directional prediction about the differences of elemental composition of protein sequences along environmental redox gradients. We found initial support for this prediction by observing depth-wise decreases of redox potential and carbon oxidation state of community reference proteomes from data reported for different Winogradsky columns. Our main analysis of 68 data sets for bacterial communities uncovered predominantly positive Z C–Eh7 correlations at local scales and a positive correlation globally in each of the seven environment types. These results suggest that chemical differences between protein sequences occurring at the millimeter scale or 1,000-km scale may, to some extent, be explained by a single hypothesis of energy minimization shaping protein evolution. Our conclusions are based on protein sequences from reference genomes and do not account for protein abundances. We could not find metaproteomic data sets that would permit comparing protein abundances with Eh measurements. Analysis of data from selected studies that reported both metaproteomes and 16S rRNA gene sequences revealed a positive but weak correlation between metaproteomic and 16S rRNA-based estimates of carbon oxidation state. The higher values for metaproteomes appear to be consistent with higher expression of relatively oxidized proteins, but we cannot rule out that the differences may be due to technical bias. Our findings imply that redox gradients may be major chemical drivers of community structure, in addition to previously recognized physicochemical factors including temperature and pH (120). Because of the negative or non-existent association between Z C and Eh7 in some individual data sets and the scatter in the global-scale analysis, we infer that geochemistry modulates, rather than controls, the underlying evolutionary and ecological processes that alter genome sequences and relative abundances of taxa. Based on these findings, it can be expected that redox potential is one of many factors that influence the elemental composition of bacterial communities, but elemental analysis of biomass, amino acid analysis of bulk protein, or measurements of metaproteomic abundances are needed to more directly test that prediction. Experimental evolution under defined redox conditions, such as the evolution of E. coli under aerobic or anaerobic conditions (121), could provide experimental tests of the chemical adaptation hypothesis. Supplementary Material Reviewer comments ACKNOWLEDGMENTS We are grateful to Matti Ruuskanen for providing mapping files to demultiplex the sequence data for Lake Hazen. D.M. acknowledges funding from the Natural Science Foundation of Changsha (kq2202089) and the Key Research and Development Program of Hunan Province (2022SK2076 and 2020WK2022). J.M.D. conceived the idea for the study. D.M. provided data. J.M.D. wrote the paper with input from D.M. Both authors participated in revising the paper and interpreting results. We declare that we have no conflicts of interest. DATA AVAILABILITY The data sets supporting the conclusions of this article are available in two R packages. chem16S is developed on GitHub (https://github.com/jedick/chem16S); version 0.1.3 of the package used in this study is archived on Zenodo (122). chem16S was used for calculating chemical metrics of community reference proteomes from RDP Classifier output and includes the amino acid compositions of reference proteomes for taxa derived from RefSeq. JMDplots is developed on GitHub (https://github.com/jedick/JMDplots); version 1.2.17 of the package used in this study is archived on Zenodo (123). JMDplots has been developed to support multiple studies; the sequence processing script, RDP Classifier output, compiled sample metadata, and code written for this study are specifically available in the orp16S section of the package, and the orp16S. Rmd vignette runs the code to make each of the figures in the paper. SUPPLEMENTAL MATERIAL The following material is available online at https://doi.org/10.1128/msystems.00014-23. 10.1128/msystems.00014-23.SuF1 FIG S1 msystems.00014-23-s0001.pdf ZC–Eh7 scatterplots and linear regressions for all data sets. Subtitles indicate the number of samples, Pearson correlation coefficient, and slope of the linear regression ± the margin of error for the 95% confidence interval. Regression lines are solid if the slope is > 0.01/V or < −0.01/V, or dashed otherwise. Click here for additional data file. 10.1128/msystems.00014-23.SuF2 FIG S2 msystems.00014-23-s0002.pdf Comparison of Eh7, Eh, and O2 concentration as predictors of carbon oxidation state. Within each domain, all plots represent the same set of samples (i.e., those for which O2 measurements are available). The number of data sets is indicated by bold numbers at the top left of each plot; the number of samples and slope of the linear regression ± the margin of error for the 95% confidence interval are shown at upper right; the Pearson correlation coefficient is shown at bottom right. The numbers of samples for each environment type are listed in the bottom legends. Click here for additional data file. 10.1128/msystems.00014-23.SuF3 FIG S3 msystems.00014-23-s0003.pdf Global analysis including only data sets generated with 515F/806R primers (these data sets are identified in Table S1). This figure was made analogously to Fig. 6 Click here for additional data file. 10.1128/msystems.00014-23.SuF4 TABLE S1 msystems.00014-23-s0004.xlsx Data set summaries: Bibliographic key, name, 515F/806R primer set (yes/no), number of samples with bacterial and archaeal sequences, sample type (soil, sediment, water, or biological), ranges of T (°C), pH, Eh (mV), Eh7 (mV), sign of Z C–Eh7 correlation for bacteria, sources of Eh and O2 data, notes. Click here for additional data file. 10.1128/msystems.00014-23.SuF5 TABLE S2 msystems.00014-23-s0005.xlsx Sequence processing statistics. Click here for additional data file. 10.1128/msystems.00014-23.SuF6 OPEN PEER REVIEW reviewer-comments.pdf An accounting of the reviewer feedback and comments. Click here for additional data file. ASM does not own the copyrights to Supplemental Material that may be linked to, or accessed through, an article. The authors have granted ASM a non-exclusive, world-wide license to publish the Supplemental Material files. Please contact the corresponding author directly for reuse. ==== Refs REFERENCES 1 Ren M , Feng X , Huang Y , Wang H , Hu Z , Clingenpeel S , Swan BK , Fonseca MM , Posada D , Stepanauskas R , Hollibaugh JT , Foster PG , Woyke T , Luo H . 2019. Phylogenomics suggests oxygen availability as a driving force in Thaumarchaeota evolution. ISME J 13 :2150–2161. doi:10.1038/s41396-019-0418-8 31024152 2 Algeo TJ , Li C . 2020. Redox classification and calibration of redox thresholds in sedimentary systems. Geochim Cosmochim Acta 287 :8–26. doi:10.1016/j.gca.2020.01.055 3 Dick JM . 2022. A thermodynamic model for water activity and redox potential in evolution and development. J Mol Evol 90 :182–199. doi:10.1007/s00239-022-10051-7 35279735 4 Pfister CA , Light SH , Bohannan B , Schmidt T , Martiny A , Hynson NA , Devkota S , David L , Whiteson K . 2022. Conceptual exchanges for understanding free-living and host-associated microbiomes. mSystems 7 : e0137421. doi:10.1128/msystems.01374-21 35014872 5 Dick JM , Tan J . 2023. Chemical links between redox conditions and estimated community Proteomes from 16S rRNA and reference protein sequences. Microb Ecol 85 :1338–1355. doi:10.1007/s00248-022-01988-9 35503575 6 Glud RN . 2008. Oxygen dynamics of marine sediments. Marine Biology Research 4 :243–289. doi:10.1080/17451000801888726 7 Lovley DR , Goodwin S . 1988. Hydrogen concentrations as an indicator of the predominant terminal electron-accepting reactions in aquatic sediments. Geochim Cosmochim Acta 52 :2993–3003. doi:10.1016/0016-7037(88)90163-9 8 Kahlert H . 2010. Reference electrodes, p 291–308. In Scholz F (ed), Electroanalytical methods. Springer. doi:10.1007/978-3-642-02915-8 9 Bricker OP . 1982. Redox potential: its measurement and importance in water systems, p 55–83. In Minear RA , LH Keith (ed), Inorganic species. Academic Press. doi:10.1016/B978-0-12-498301-4.50007-4 10 Dick JM , LaRowe DE , Helgeson HC . 2006. Temperature, pressure, and electrochemical constraints on protein speciation: group additivity calculation of the standard molal thermodynamic properties of ionized unfolded proteins. Biogeosciences 3 :311–336. doi:10.5194/bg-3-311-2006 11 Dick JM . 2019. CHNOSZ: Thermodynamic calculations and diagrams for geochemistry. Front Earth Sci 7 :180. doi:10.3389/feart.2019.00180 12 Dick JM , Yu M , Tan J . 2020. Uncovering chemical signatures of salinity gradients through compositional analysis of protein sequences. Biogeosciences 17 :6145–6162. doi:10.5194/bg-17-6145-2020 13 Leinonen R , Sugawara H , Shumway M , International Nucleotide Sequence Database Collaboration . 2011. The sequence read archive. Nucleic Acids Res 39 :D19–D21. doi:10.1093/nar/gkq1019 21062823 14 Ghannam RB , Schaerer LG , Butler TM , Techtmann SM . 2020. Biogeographic patterns in members of globally distributed and dominant taxa found in port microbial communities. mSphere 5 : e00481-19. doi:10.1128/mSphere.00481-19 31996419 15 Power JF , Carere CR , Lee CK , Wakerley GLJ , Evans DW , Button M , White D , Climo MD , Hinze AM , Morgan XC , McDonald IR , Cary SC , Stott MB . 2018. Microbial biogeography of 925 geothermal springs in New Zealand. Nat Commun 9 :2876. doi:10.1038/s41467-018-05020-y 30038374 16 Rundell EA , Banta LM , Ward DV , Watts CD , Birren B , Esteban DJ . 2014. 16S rRNA gene survey of microbial communities in Winogradsky columns. PLoS One 9 : e104134. doi:10.1371/journal.pone.0104134 25101630 17 Diez-Ercilla M , Falagán C , Yusta I , Sánchez-España J . 2019. Metal mobility and mineral transformations driven by bacterial activity in acidic pit lake sediments: evidence from column experiments and sequential extraction. J Soils Sediments 19 :1527–1542. doi:10.1007/s11368-018-2112-2 18 Mai Y-Z , Lai Z-N , Li X-H , Peng S-Y , Wang C . 2018. Structural and functional shifts of bacterioplanktonic communities associated with spatiotemporal gradients in river outlets of the subtropical Pearl River Estuary, South China. Mar Pollut Bull 136 :309–321. doi:10.1016/j.marpolbul.2018.09.013 30509812 19 He P , Xie L , Zhang X , Li J , Lin X , Pu X , Yuan C , Tian Z , Li J . 2020. Microbial diversity and metabolic potential in the stratified Sansha Yongle Blue Hole in the South China Sea. Sci Rep 10 :5949. doi:10.1038/s41598-020-62411-2 32249806 20 Xie L , Wang B , Pu X , Xin M , He P , Li C , Wei Q , Zhang X , Li T . 2019. Hydrochemical properties and chemocline of the Sansha Yongle Blue Hole in the South China Sea. Sci Total Environ 649 :1281–1292. doi:10.1016/j.scitotenv.2018.08.333 30308898 21 Wang L , Han M , Li X , Yu B , Wang H , Ginawi A , Ning K , Yan Y . 2021. Mechanisms of niche-neutrality balancing can drive the assembling of microbial community. Mol Ecol 30 :1492–1504. doi:10.1111/mec.15825 33522045 22 Zhang S , Li K , Hu J , Wang F , Chen D , Zhang Z , Li T , Li L , Tao J , Liu D , Che R . 2022. Distinct assembly mechanisms of microbial sub-communities with different rarity along the Nu River. J Soils Sediments 22 :1530–1545. doi:10.1007/s11368-022-03149-4 23 Zhang L , Zhang C , Lian K , Ke D , Xie T , Liu C . 2021. River restoration changes distributions of antibiotics, antibiotic resistance genes, and microbial community. Sci Total Environ 788 :147873. doi:10.1016/j.scitotenv.2021.147873 34134371 24 Li C , Wang L , Ji S , Chang M , Wang L , Gan Y , Liu J . 2021. The ecology of the plastisphere: Microbial composition, function, assembly, and network in the freshwater and seawater ecosystems. Water Res 202 : 117428. doi:10.1016/j.watres.2021.117428 34303166 25 Gao Y , Zhang W , Li Y . 2021. Microbial community coalescence: Does it matter in the Three Gorges Reservoir? Water Res 205 :117428. doi:10.1016/j.watres.2021.117638 26 Ramos-Perez D , Alcántara-Hernández RJ , Romero FM , González-Chávez JL . 2022. Changes in the prokaryotic diversity in response to hydrochemical variations during an acid mine drainage passive treatment. Sci Total Environ 842 : 156629. doi:10.1016/j.scitotenv.2022.156629 35691343 27 Schneider D , Arp G , Reimer A , Reitner J , Daniel R . 2013. Phylogenetic analysis of a microbialite-forming microbial mat from a hypersaline lake of the Kiritimati Atoll, Central Pacific. PLoS One 8 : e66662. doi:10.1371/journal.pone.0066662 23762495 28 Liu M , Liu L , Chen H , Yu Z , Yang JR , Xue Y , Huang B , Yang J . 2019. Community dynamics of free-living and particle-attached bacteria following a reservoir Microcystis bloom. Sci Total Environ 660 :501–511. doi:10.1016/j.scitotenv.2018.12.414 30640117 29 Baricz A , Chiriac CM , Andrei A-Ștefan , Bulzu P-A , Levei EA , Cadar O , Battes KP , Cîmpean M , Șenilă M , Cristea A , Muntean V , Alexe M , Coman C , Szekeres EK , Sicora CI , Ionescu A , Blain D , O’Neill WK , Edwards J , Hallsworth JE , Banciu HL . 2021. Spatio-temporal insights into microbiology of the freshwater-to-hypersaline, oxic-hypoxic-euxinic waters of Ursu Lake. Environ Microbiol 23 :3523–3540. doi:10.1111/1462-2920.14909 31894632 30 Hu A , Li S , Zhang L , Wang H , Yang J , Luo Z , Rashid A , Chen S , Huang W , Yu C-P . 2018. Prokaryotic footprints in urban water ecosystems: A case study of urban landscape ponds in a coastal city, China. Environ Pollut 242 :1729–1739. doi:10.1016/j.envpol.2018.07.097 30064876 31 Butler TM , Wilhelm A-C , Dwyer AC , Webb PN , Baldwin AL , Techtmann SM . 2019. Microbial community dynamics during lake ice freezing. Sci Rep 9 : 6231. doi:10.1038/s41598-019-42609-9 30996247 32 Ionescu D , Bizic M , Karnatak R , Musseau CL , Onandia G , Kasada M , Berger SA , Nejstgaard JC , Ryo M , Lischeid G , Gessner MO , Wollrab S , Grossart H‐. P . 2022. From microbes to mammals: Pond biodiversity homogenization across different land‐use types in an agricultural landscape. Ecological Monographs 92 : 3. doi:10.1002/ecm.1523 33 Ninio S , Lupu A , Eckert W , Ostrovsky I , Viner Mozzini Y , Sukenik A . 2021. Metalimnetic chlorophyll maxima in Lake Kinneret ‐ Chlorobium revisited. Freshw Biol 66 :468–480. doi:10.1111/fwb.13653 34 Menéndez-Serra M , Triadó-Margarit X , Casamayor EO . 2021. Ecological and metabolic thresholds in the bacterial, protist, and fungal microbiome of ephemeral saline lakes (Monegros Desert, Spain). Microb Ecol 82 :885–896. doi:10.1007/s00248-021-01732-9 33725151 35 Sanseverino I , Pretto P , António DC , Lahm A , Facca C , Loos R , Skejo H , Beghi A , Pandolfi F , Genoni P , Lettieri T . 2022. Metagenomics analysis to investigate the microbial communities and their functional profile during cyanobacterial blooms in Lake Varese. Microb Ecol 83 :850–868. doi:10.1007/s00248-021-01914-5 34766210 36 Guo L , Wang G , Sheng Y , Sun X , Shi Z , Xu Q , Mu W . 2020. Temperature governs the distribution of hot spring microbial community in three hydrothermal fields, Eastern Tibetan Plateau Geothermal Belt, Western China. Sci Total Environ 720 : 137574. doi:10.1016/j.scitotenv.2020.137574 32145630 37 Peltek SE , Bryanskaya AV , Uvarova YE , Rozanov AS , Ivanisenko TV , Ivanisenko VA , Lazareva EV , Saik OV , Efimov VM , Zhmodik SM , Taran OP , Slynko NM , Shekhovtsov SV , Parmon VN , Dobretsov NL , Kolchanov NA . 2020. Young «oil site» of the Uzon Caldera as a habitat for unique microbial life. BMC Microbiol 20 :349. doi:10.1186/s12866-020-02012-1 33228530 38 Ma L , Wu G , Yang J , Huang L , Phurbu D , Li W-J , Jiang H . 2021. Distribution of hydrogen-producing bacteria in Tibetan hot springs, China. Front Microbiol 12 :569020. doi:10.3389/fmicb.2021.569020 34367076 39 Sabuda MC , Brazelton WJ , Putman LI , McCollom TM , Hoehler TM , Kubo MDY , Cardace D , Schrenk MO . 2020. A dynamic microbial sulfur cycle in a serpentinizing continental ophiolite. Environ Microbiol 22 :2329–2345. doi:10.1111/1462-2920.15006 32249550 40 Rempfert KR , Miller HM , Bompard N , Nothaft D , Matter JM , Kelemen P , Fierer N , Templeton AS . 2017. Geological and geochemical controls on subsurface microbial life in the Samail Ophiolite, Oman. Front Microbiol 8 :56. doi:10.3389/fmicb.2017.00056 28223966 41 Crespo-Medina M , Twing KI , Sánchez-Murillo R , Brazelton WJ , McCollom TM , Schrenk MO . 2017. Methane dynamics in a tropical serpentinizing environment: The Santa Elena Ophiolite, Costa Rica. Front Microbiol 8 :916. doi:10.3389/fmicb.2017.00916 28588569 42 Kamran A , Sauter K , Reimer A , Wacker T , Reitner J , Hoppert M . 2020. Cyanobacterial mats in calcite-precipitating serpentinite-hosted alkaline springs of the Voltri Massif, Italy. Microorganisms 9 : 62. doi:10.3390/microorganisms9010062 33383678 43 Putman LI , Sabuda MC , Brazelton WJ , Kubo MD , Hoehler TM , McCollom TM , Cardace D , Schrenk MO . 2021. Microbial communities in a serpentinizing aquifer are assembled through strong concurrent dispersal limitation and selection. mSystems 6 : e0030021. doi:10.1128/mSystems.00300-21 34519519 44 Nothaft DB , Templeton AS , Boyd ES , Matter JM , Stute M , Paukert Vankeuren AN , The Oman Drilling Project Science Team . 2021. Aqueous geochemical and microbial variation across discrete depth intervals in a peridotite aquifer assessed using a packer system in the Samail Ophiolite, Oman. JGR Biogeosciences 126 : e2021JG006319. doi:10.1029/2021JG006319 45 Kocur CMD , Lomheim L , Molenda O , Weber KP , Austrins LM , Sleep BE , Boparai HK , Edwards EA , O’Carroll DM . 2016. Long-term field study of microbial community and dechlorinating activity following carboxymethyl cellulose-stabilized nanoscale zero-valent iron injection. Environ Sci Technol 50 :7658–7670. doi:10.1021/acs.est.6b01745 27305345 46 Wang Y , Li P , Jiang Z , Sinkkonen A , Wang S , Tu J , Wei D , Dong H , Wang Y . 2016. Microbial community of high arsenic groundwater in agricultural irrigation area of Hetao Plain, Inner Mongolia. Front Microbiol 7 :1917. doi:10.3389/fmicb.2016.01917 27999565 47 Zheng T , Deng Y , Wang Y , Jiang H , O’Loughlin EJ , Flynn TM , Gan Y , Ma T . 2019. Seasonal microbial variation accounts for arsenic dynamics in shallow alluvial aquifer systems. J Hazard Mater 367 :109–119. doi:10.1016/j.jhazmat.2018.12.087 30594709 48 Danczak RE , Johnston MD , Kenah C , Slattery M , Wilkins MJ . 2018. Microbial community cohesion mediates community turnover in unperturbed aquifers. mSystems 3 : e00066-18. doi:10.1128/mSystems.00066-18 29984314 49 Sonthiphand P , Ruangroengkulrith S , Mhuantong W , Charoensawan V , Chotpantarat S , Boonkaewwan S . 2019. Metagenomic insights into microbial diversity in a groundwater basin impacted by a variety of anthropogenic activities. Environ Sci Pollut Res Int 26 :26765. doi:10.1007/s11356-019-05905-5 31300992 50 Aguilar-Rangel EJ , Prado BL , Vásquez-Murrieta MS , Los Santos PE , Siebe C , Falcón LI , Santillán J , Alcántara-Hernández RJ . 2020. Temporal analysis of the microbial communities in a nitrate-contaminated aquifer and the co-occurrence of anammox, n-damo and nitrous-oxide reducing bacteria. J Contam Hydrol 234 :103657. doi:10.1016/j.jconhyd.2020.103657 32777591 51 Yan L , Herrmann M , Kampe B , Lehmann R , Totsche KU , Küsel K . 2020. Environmental selection shapes the formation of near-surface groundwater microbiomes. Water Res 170 :115341. doi:10.1016/j.watres.2019.115341 31790889 52 Zecchin S , Crognale S , Zaccheo P , Fazi S , Amalfitano S , Casentini B , Callegari M , Zanchi R , Sacchi GA , Rossetti S , Cavalca L . 2021. Adaptation of microbial communities to environmental arsenic and selection of arsenite-oxidizing bacteria from contaminated groundwaters. Front Microbiol 12 :634025. doi:10.3389/fmicb.2021.634025 33815317 53 Mosley OE , Gios E , Weaver L , Close M , Daughney C , van der Raaij R , Martindale H , Handley KM . 2022. Metabolic diversity and aero-tolerance in anammox bacteria from geochemically distinct aquifers. mSystems 7 : e0125521. doi:10.1128/msystems.01255-21 35191775 54 Morrissy JG , Currell MJ , Reichman SM , Surapaneni A , Megharaj M , Crosbie ND , Hirth D , Aquilina S , Rajendram W , Ball AS . 2022. The variation in groundwater microbial communities in an unconfined aquifer contaminated by multiple nitrogen contamination sources. Water 14 :613. doi:10.3390/w14040613 55 Zhou Z , Meng H , Liu Y , Gu J-D , Li M . 2017. Stratified bacterial and archaeal community in mangrove and intertidal wetland mudflats revealed by high throughput 16S rRNA gene sequencing. Front Microbiol 8 :2148. doi:10.3389/fmicb.2017.02148 29163432 56 Broman E , Sjöstedt J , Pinhassi J , Dopson M . 2017. Shifts in coastal sediment oxygenation cause pronounced changes in microbial community composition and associated metabolism. Microbiome 5 :96. doi:10.1186/s40168-017-0311-5 28793929 57 Rissanen AJ , Karvinen A , Nykänen H , Peura S , Tiirola M , Mäki A , Kankaala P . 2017. Effects of alternative electron acceptors on the activity and community structure of methane-producing and consuming microbes in the sediments of two shallow boreal lakes. FEMS Microbiol Ecol 93 : fix078. doi:10.1093/femsec/fix078 58 Han M , Dsouza M , Zhou C , Li H , Zhang J , Chen C , Yao Q , Zhong C , Zhou H , Gilbert JA , Wang Z , Ning K . 2019. Agricultural risk factors influence microbial ecology in Honghu Lake. Genomics Proteomics Bioinformatics 17 :76–90. doi:10.1016/j.gpb.2018.04.008 31026580 59 Otte JM , Harter J , Laufer K , Blackwell N , Straub D , Kappler A , Kleindienst S . 2018. The distribution of active iron-cycling bacteria in marine and freshwater sediments is decoupled from geochemical gradients. Environ Microbiol 20 :2483–2499. doi:10.1111/1462-2920.14260 29708639 60 Wu J , Hong Y , Liu X , Hu Y . 2021. Variations in nitrogen removal rates and microbial communities over sediment depth in Daya Bay, China. Environ Pollut 286 :117267. doi:10.1016/j.envpol.2021.117267 33965803 61 Ruuskanen MO , St Pierre KA , St Louis VL , Aris-Brosou S , Poulain AJ . 2018. Physicochemical drivers of microbial community structure in sediments of Lake Hazen, Nunavut, Canada. Front Microbiol 9 :1138. doi:10.3389/fmicb.2018.01138 29922252 62 Cai W , Li Y , Shen Y , Wang C , Wang P , Wang L , Niu L , Zhang W . 2019. Vertical distribution and assemblages of microbial communities and their potential effects on sulfur metabolism in a black-odor urban river. J Environ Manage 235 :368–376. doi:10.1016/j.jenvman.2019.01.078 30708274 63 von Hoyningen-Huene AJE , Schneider D , Fussmann D , Reimer A , Arp G , Daniel R . 2019. Bacterial succession along a sediment porewater gradient at Lake Neusiedl in Austria. Sci Data 6 :163. doi:10.1038/s41597-019-0172-9 31471542 64 Zhang K , Hu Z , Zeng F , Yang X , Wang J , Jing R , Zhang H , Li Y , Zhang Z . 2019. Biodegradation of petroleum hydrocarbons and changes in microbial community structure in sediment under nitrate-, ferric-, sulfate-reducing and methanogenic conditions. J Environ Manage 249 :109425. doi:10.1016/j.jenvman.2019.109425 31446121 65 Lanzén A , Mendibil I , Borja Á , Alonso-Sáez L . 2021. A microbial mandala for environmental monitoring: Predicting multiple impacts on estuarine prokaryote communities of the Bay of Biscay. Mol Ecol 30 :2969–2987. doi:10.1111/mec.15489 32479653 66 AJE von , Schneider D , Fussmann D , Reimer A , Arp G , Daniel R . 2022. DNA- and RNA-based bacterial communities and geochemical zonation under changing sediment porewater dynamics on the Aldabra Atoll. Sci Rep 12 :4257. doi:10.1038/s41598-022-07980-0 35277525 67 Zhang L , Zhang C , Lian K , Liu C . 2021. Effects of chronic exposure of antibiotics on microbial community structure and functions in hyporheic zone sediments. J Hazard Mater 416 : 126141. doi:10.1016/j.jhazmat.2021.126141 34492930 68 Wyness AJ , Fortune I , Blight AJ , Browne P , Hartley M , Holden M , Paterson DM . 2021. Ecosystem engineers drive differing microbial community composition in intertidal estuarine sediments. PLoS One 16 : e0240952. doi:10.1371/journal.pone.0240952 33606695 69 Hu Y , Chen J , Wang C , Wang P , Gao H , Zhang J , Zhang B , Cui G , Zhao D . 2022. Insight into microbial degradation of hexabromocyclododecane (HBCD) in lake sediments under different hydrodynamic conditions. Sci Total Environ 827 : 154358. doi:10.1016/j.scitotenv.2022.154358 35259383 70 Wolińska A , Kruczyńska A , Grządziel J , Gałązka A , Marzec-Grządziel A , Szałaj K , Kuźniar A . 2022. Functional and seasonal changes in the structure of microbiome inhabiting bottom sediments of a pond intended for ecological king carp farming. Biology (Basel) 11 : 913. doi:10.3390/biology11060913 35741434 71 Meng D , Li J , Liu T , Liu Y , Yan M , Hu J , Li X , Liu X , Liang Y , Liu H , Yin H . 2019. Effects of redox potential on soil cadmium solubility: Insight into microbial community. J Environ Sci (China) 75 :224–232. doi:10.1016/j.jes.2018.03.032 30473288 72 Brigham BA , Montero AD , O’Mullan GD , Bird JA . 2018. Acetate additions stimulate CO2 and CH4 production from urban wetland soils. Soil Sci Soc Am J 82 :1147–1159. doi:10.2136/sssaj2018.01.0034 73 Wang Q , Han Y , Lan S , Hu C . 2021. Metagenomic insight into patterns and mechanism of nitrogen cycle during biocrust succession. Front Microbiol 12 :633428. doi:10.3389/fmicb.2021.633428 33815315 74 Chuang S , Wang B , Chen K , Jia W , Qiao W , Ling W , Tang X , Jiang J . 2020. Microbial catabolism of lindane in distinct layers of acidic paddy soils combinedly affected by different water managements and bioremediation strategies. Sci Total Environ 746 : 140992. doi:10.1016/j.scitotenv.2020.140992 32745849 75 Poret-Peterson AT , Sayed N , Glyzewski N , Forbes H , González-Orta ET , Kluepfel DA . 2020. Temporal responses of microbial communities to anaerobic soil disinfestation. Microb Ecol 80 :191–201. doi:10.1007/s00248-019-01477-6 31873773 76 Luan L , Jiang Y , Cheng M , Dini-Andreote F , Sui Y , Xu Q , Geisen S , Sun B . 2020. Organism body size structures the soil microbial and nematode community assembly at a continental and global scale. Nat Commun 11 :6406. doi:10.1038/s41467-020-20271-4 33335105 77 Dai J , Tang Z , Jiang N , Kopittke PM , Zhao F-J , Wang P . 2020. Increased arsenic mobilization in the rice rhizosphere is mediated by iron-reducing bacteria. Environ Pollut 263 : 114561. doi:10.1016/j.envpol.2020.114561 32320889 78 Rolando JL , Kolton M , Song T , Kostka JE . 2022. The core root microbiome of Spartina alterniflora is predominated by sulfur-oxidizing and sulfate-reducing bacteria in Georgia salt marshes, USA. Microbiome 10 :37. doi:10.1186/s40168-021-01187-7 35227326 79 Dykes GE , Limmer MA , Seyfferth AL . 2021. Silicon-rich soil amendments impact microbial community composition and the composition of arsM bearing microbes. Plant Soil 468 :147–164. doi:10.1007/s11104-021-05103-8 80 Limmer MA , Seyfferth AL . 2021. Carryover effects of silicon‐rich amendments in rice paddies. Soil Sci. Soc. Am. J 85 :314–327. doi:10.1002/saj2.20146 81 Wolińska A , Kruczyńska A , Podlewski J , Słomczewski A , Grządziel J , Gałązka A , Kuźniar A . 2022. Does the use of an intercropping mixture really improve the biology of monocultural soils?—A search for bacterial indicators of sensitivity and resistance to long-term maize monoculture. Agronomy 12 :613. doi:10.3390/agronomy12030613 82 Chowdhury SA , Kaneko A , Baki MZI , Takasugi C , Wada N , Asiloglu R , Harada N , Suzuki K . 2022. Impact of the chemical composition of applied organic materials on bacterial and archaeal community compositions in paddy soil. Biol Fertil Soils 58 :135–148. doi:10.1007/s00374-022-01619-y 83 Cai YJ , Liu ZA , Zhang S , Liu H , Nicol GW , Chen Z . 2022. Microbial community structure is stratified at the millimeter-scale across the soil–water interface. ISME Commun 2 :53. doi:10.1038/s43705-022-00138-z 84 Meier DV , Pjevac P , Bach W , Markert S , Schweder T , Jamieson J , Petersen S , Amann R , Meyerdierks A . 2019. Microbial metal-sulfide oxidation in inactive hydrothermal vent chimneys suggested by metagenomic and metaproteomic analyses. Environ Microbiol 21 :682–701. doi:10.1111/1462-2920.14514 30585382 85 Reeves EP , Yoshinaga MY , Pjevac P , Goldenstein NI , Peplies J , Meyerdierks A , Amann R , Bach W , Hinrichs K-U . 2014. Microbial lipids reveal carbon assimilation patterns on hydrothermal sulfide chimneys. Environ Microbiol 16 :3515–3532. doi:10.1111/1462-2920.12525 24905086 86 Kleiner M , Thorson E , Sharp CE , Dong X , Liu D , Li C , Strous M . 2017. Assessing species biomass contributions in microbial communities via metaproteomics. Nat Commun 8 :1558. doi:10.1038/s41467-017-01544-x 29146960 87 Hawley AK , Torres-Beltrán M , Zaikova E , Walsh DA , Mueller A , Scofield M , Kheirandish S , Payne C , Pakhomova L , Bhatia M , Shevchuk O , Gies EA , Fairley D , Malfatti SA , Norbeck AD , Brewer HM , Pasa-Tolic L , Del Rio TG , Suttle CA , Tringe S , Hallam SJ . 2017. A compendium of multi-omic sequence information from the Saanich Inlet water column. Sci Data 4 :170160. doi:10.1038/sdata.2017.160 29087368 88 Husson O , Husson B , Brunet A , Babre D , Alary K , Sarthou J-P , Charpentier H , Durand M , Benada J , Henry M . 2016. Practical improvements in soil redox potential (Eh) measurement for characterisation of soil properties. Application for comparison of conventional and conservation agriculture cropping systems. Anal Chim Acta 906 :98–109. doi:10.1016/j.aca.2015.11.052 26772129 89 ZoBell CE . 1946. Studies on redox potential of marine sediments. Am Assoc Pet Geol Bull 30 :477–513. doi:10.1306/3D933808-16B1-11D7-8645000102C1865D 90 Becking LGMB , Kaplan IR , Moore D . 1960. Limits of the natural environment in terms of pH and oxidation-reduction potentials. J Geol 68 :243–284. doi:10.1086/626659 91 Rognes T , Flouri T , Nichols B , Quince C , Mahé F . 2016. VSEARCH: A versatile open source tool for metagenomics. PeerJ 4 : e2584. doi:10.7717/peerj.2584 27781170 92 Quast C , Pruesse E , Yilmaz P , Gerken J , Schweer T , Yarza P , Peplies J , Glöckner FO . 2013. The SILVA ribosomal RNA gene database project: Improved data processing and web-based tools. Nucleic Acids Res 41 :D590–D596. doi:10.1093/nar/gks1219 23193283 93 Cole JR , Wang Q , Fish JA , Chai B , McGarrell DM , Sun Y , Brown CT , Porras-Alfaro A , Kuske CR , Tiedje JM . 2014. Ribosomal Database Project: Data and tools for high throughput rRNA analysis. Nucleic Acids Res 42 :D633–D642. doi:10.1093/nar/gkt1244 24288368 94 Li W , O’Neill KR , Haft DH , DiCuccio M , Chetvernin V , Badretdin A , Coulouris G , Chitsaz F , Derbyshire MK , Durkin AS , Gonzales NR , Gwadz M , Lanczycki CJ , Song JS , Thanki N , Wang J , Yamashita RA , Yang M , Zheng C , Marchler-Bauer A , Thibaud-Nissen F . 2021. RefSeq: Expanding the Prokaryotic Genome Annotation Pipeline reach with protein family model curation. Nucleic Acids Res 49 :D1020–D1028. doi:10.1093/nar/gkaa1105 33270901 95 Pjevac P , Meier DV , Markert S , Hentschker C , Schweder T , Becher D , Gruber-Vodicka HR , Richter M , Bach W , Amann R , Meyerdierks A . 2018. Metaproteogenomic profiling of microbial communities colonizing actively venting hydrothermal chimneys. Front Microbiol 9 :680. doi:10.3389/fmicb.2018.00680 29696004 96 Core Team R . 2022. R: a language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. Available from: https://www.R-project.org 97 Kelley D , Richards C . 2022. oce: analysis of oceanographic Data. R package version 1.7-2. Available from: https://cran.r-project.org/package=oce 98 U.S. Geological Survey . 2010. Great Lakes and Watersheds Shapefiles. https://www.sciencebase.gov/catalog/item/530f8a0ee4b0e7e46bd300dd 99 Douglas GM , Maffei VJ , Zaneveld JR , Yurgel SN , Brown JR , Taylor CM , Huttenhower C , Langille MGI . 2020. PICRUSt2 for prediction of metagenome functions. Nat Biotechnol 38 :685–688. doi:10.1038/s41587-020-0548-6 32483366 100 Anderson GM , Crerar DA . 1993. Thermodynamics in geochemistry: the equilibrium model. Oxford University Press, New York. doi:10.1093/oso/9780195064643.001.0001 101 Dick JM , Boyer GM , Canovas PA , Shock EL . 2023. Using thermodynamics to obtain geochemical information from genomes. Geobiology 21 :262–273. doi:10.1111/gbi.12532 36376996 102 Foerstner KU , von Mering C , Hooper SD , Bork P . 2005. Environments shape the nucleotide composition of genomes. EMBO Rep 6 :1208–1213. doi:10.1038/sj.embor.7400538 16200051 103 Smith TP , Mombrikotb S , Ransome E , Kontopoulos D-G , Pawar S , Bell T . 2022. Latent functional diversity may accelerate microbial community responses to temperature fluctuations. eLife 11 : e80867. doi:10.7554/eLife.80867 36444646 104 Million M , Raoult D . 2018. Linking gut redox to human microbiome. Human Microbiome Journal 10 :27–32. doi:10.1016/j.humic.2018.07.002 105 Masuda T , Saito N , Tomita M , Ishihama Y . 2009. Unbiased quantitation of Escherichia coli membrane proteome using phase transfer surfactants. Mol Cell Proteomics 8 :2770–2777. doi:10.1074/mcp.M900240-MCP200 19767571 106 Savas JN , Stein BD , Wu CC , Yates JR . 2011. Mass spectrometry accelerates membrane protein analysis. Trends Biochem Sci 36 :388–396. doi:10.1016/j.tibs.2011.04.005 21616670 107 Dick JM . 2014. Average oxidation state of carbon in proteins. J R Soc Interface 11 :20131095. doi:10.1098/rsif.2013.1095 25165594 108 Pibernat IV , Abellà CA . 1993. Evolution of the physico-chemical and biological parameters in a modified Winogradsky column. Sci Gerundensis 19 :35–46. https://hdl.handle.net/10256/5349. 109 Pagaling E , Strathdee F , Spears BM , Cates ME , Allen RJ , Free A . 2014. Community history affects the predictability of microbial ecosystem development. ISME J 8 :19–30. doi:10.1038/ismej.2013.150 23985743 110 Widder S , Allen RJ , Pfeiffer T , Curtis TP , Wiuf C , Sloan WT , Cordero OX , Brown SP , Momeni B , Shou W , Kettle H , Flint HJ , Haas AF , Laroche B , Kreft J-U , Rainey PB , Freilich S , Schuster S , Milferstedt K , van der Meer JR , Groβkopf T , Huisman J , Free A , Picioreanu C , Quince C , Klapper I , Labarthe S , Smets BF , Wang H , Isaac Newton Institute Fellows, Soyer OS . 2016. Challenges in microbial ecology: Building predictive understanding of community function and dynamics. ISME J 10 :2557–2568. doi:10.1038/ismej.2016.45 27022995 111 Pagaling E , Vassileva K , Mills CG , Bush T , Blythe RA , Schwarz-Linek J , Strathdee F , Allen RJ , Free A . 2017. Assembly of microbial communities in replicate nutrient-cycling model ecosystems follows divergent trajectories, leading to alternate stable states. Environ Microbiol 19 :3374–3386. doi:10.1111/1462-2920.13849 28677203 112 Parada AE , Needham DM , Fuhrman JA . 2016. Every base matters: Assessing small subunit rRNA primers for marine microbiomes with mock communities, time series and global field samples. Environ Microbiol 18 :1403–1414. doi:10.1111/1462-2920.13023 26271760 113 Caporaso JG , Lauber CL , Walters WA , Berg-Lyons D , Huntley J , Fierer N , Owens SM , Betley J , Fraser L , Bauer M , Gormley N , Gilbert JA , Smith G , Knight R . 2012. Ultra-high-throughput microbial community analysis on the Illumina HiSeq and MiSeq platforms. ISME J 6 :1621–1624. doi:10.1038/ismej.2012.8 22402401 114 Bédard C , Cisneros AF , Jordan D , Landry CR . 2022. Correlation between protein abundance and sequence conservation: What do recent experiments say? Curr Opin Genet Dev 77 :101984. doi:10.1016/j.gde.2022.101984 36162152 115 McLaren MR , Willis AD , Callahan BJ . 2019. Consistent and correctable bias in metagenomic sequencing experiments. eLife 8 : e46923. doi:10.7554/eLife.46923 31502536 116 Leary DH , Hervey WJ , Deschamps JR , Kusterbeck AW , Vora GJ . 2013. Which metaproteome? The impact of protein extraction bias on metaproteomic analyses. Mol Cell Probes 27 :193–199. doi:10.1016/j.mcp.2013.06.003 23831146 117 Lozupone CA , Knight R . 2007. Global patterns in bacterial diversity. Proc Natl Acad Sci U S A 104 :11436–11440. doi:10.1073/pnas.0611525104 17592124 118 Ortiz-Alvarez R , Casamayor EO . 2016. High occurrence of Pacearchaeota and Woesearchaeota (Archaea superphylum DPANN) in the surface waters of oligotrophic high-altitude lakes. Environ Microbiol Rep 8 :210–217. doi:10.1111/1758-2229.12370 26711582 119 Gelsinger DR , Dallon E , Reddy R , Mohammad F , Buskirk AR , DiRuggiero J . 2020. Ribosome profiling in archaea reveals leaderless translation, novel translational initiation sites, and ribosome pausing at single codon resolution. Nucleic Acids Res 48 :5201–5216. doi:10.1093/nar/gkaa304 32382758 120 Xu X , Wang N , Lipson D , Sinsabaugh R , Schimel J , He L , Soudzilovskaia NA , Tedersoo L , Algar AC . 2020. Microbial macroecology: In search of mechanisms governing microbial biogeographic patterns. Global Ecol. Biogeogr 29 :1870–1886. doi:10.1111/geb.13162 121 Finn TJ , Shewaramani S , Leahy SC , Janssen PH , Moon CD . 2017. Dynamics and genetic diversification of Escherichia coli during experimental adaptation to an anaerobic environment. PeerJ 5 : e3244. doi:10.7717/peerj.3244 28480139 122 Dick JM . 2023. chem16S 0.13. Zenodo. doi:10.5281/zenodo.7633130 123 Dick JM . 2023. JMDplots 1.2.17. Zenodo. doi:10.5281/zenodo.7634341