
==== Front
Am J Respir Crit Care Med
Am J Respir Crit Care Med
ajrccm
American Journal of Respiratory and Critical Care Medicine
1073-449X
1535-4970
American Thoracic Society

202404-0835RL
10.1164/rccm.202404-0835RL
Correspondence
Approaches for Mycobacterium tuberculosis Transmission Inference Based on Genomic Data
Cohen Ted 1
Colijn Caroline 3
Warren Joshua L. 2
1 Department of Epidemiology of Microbial Diseases and
2 Department of Biostatistics, Yale School of Public Health, New Haven, Connecticut; and
3 Department of Mathematics, Simon Fraser University, Burnaby, British Columbia, Canada
Correspondence and requests for reprints should be addressed to Ted Cohen, M.D., D.P.H., Department of Epidemiology of Microbial Diseases, Yale School of Public Health, 60 College Street, New Haven, CT 06510. Email: theodore.cohen@yale.edu.
17 7 2024
15 9 2024
17 7 2024
210 6 847849
Copyright © 2024 by the American Thoracic Society
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This article is open access and distributed under the terms of the Creative Commons Attribution Non-Commercial No Derivatives License 4.0. For commercial usage and reprints, please e-mail Diane Gern (dgern@thoracic.org).
==== Body
pmcTo the Editor:

Advances in genomic sequencing and the development of new analytic methods have enabled the detailed reconstruction of Mycobacterium tuberculosis (and other pathogen) outbreaks. These technologies and methods create new opportunities to identify factors associated with pathogen transmission between individuals. In a recent study reported in the Journal, Trevisi and colleagues (1) used whole-genome sequencing to identify transmission pairs and the directionality of transmission of M. tuberculosis in Lima, Peru. However, we have several concerns about how transmission pairs were identified and how factors associated with transmission were estimated.

A common analytic approach in this setting is to consider pairs (dyads) of infected individuals and estimate factors associated with these individuals’ being linked through transmission on the basis of the pairwise genetic similarity of pathogen sequences. In Trevisi and colleagues’ study (1), each putative M. tuberculosis transmission pair was treated as an independent observation during analysis. However, network data such as these often exhibit potentially complex correlation patterns (2), and working in the infectious disease setting, we previously demonstrated that failure to account for such correlation can compromise the validity of such analyses (3) and developed an R (https://www.r-project.org) package, GenePair, to facilitate the implementation of these approaches (https://github.com/warrenjl/GenePair). We note that several other frameworks have also been developed for properly analyzing data that are correlated because of their network structure (see Reference 4 for a recent review).

These data are likely to be correlated because a single infected individual is present across multiple dyads, leading to dependencies between the outcomes describing the genetic similarity of pathogens infecting individuals (e.g., binary classifications of transmission/no transmission, as was used by Trevisi and colleagues [1], or continuous outcomes such as the number of SNPs or transmission probabilities). For example, if individual A has a strain genetically similar to those infecting individuals B and C, then pathogens from B and C are much more likely to be genetically similar than from two randomly chosen individuals; indeed, there is a maximum distance B and C could possibly be from each other. Outcomes across the A–B, A–C, and B–C pairs should not be treated as independent observations, because B and C must be close if both A–B and B–C are. This problem expands if pathogens from A, B, and C are all genetically close to many others, for example, in a large, closely related cluster. The number of pairs (assuming symmetry in the outcome) in a cluster of size n is n choose 2: 190 pairs for 20 individuals, 1,225 pairs for 50 individuals, and 4,950 pairs for 100 individuals. In a cluster with 50 cases (and 1 index case), there are 49 true transmission events but 1,225 pairs. Statistical analyses of these paired outcomes that ignore the network dependencies can underestimate the uncertainty in key associations of interest, resulting in an increased risk of type I errors (5). Through simulation, we also found that point estimation could be adversely affected in these cases (3).

Here, we demonstrate the utility of GenePair using previously published tuberculosis data from Moldova as an example (6). To parallel the analytic approach presented by Trevisi and colleagues (1), we fit a logistic regression model with an outcome of whether two individuals are in a transmission pair (as defined by a SNP difference of three or less) and with several predictors: whether the individuals come from the same village, the distance between home villages, the time between diagnosis dates, and the age difference of the pair. We analyzed the data accounting for correlation across pairs (using GenePair) and with a standard logistic regression model (i.e., without accounting for correlations across pairs). We also analyzed the data using GenePair after randomly shuffling the pair labels (i.e., eliminating the correlation due to the paired structure of the data) to determine the impact of applying these methods when they are not needed.

Figure 1 shows posterior inference (i.e., posterior means and 95% credible intervals) for all fitted models. As expected, in the presence of correlation, GenePair resulted in wider credible intervals than the standard approach and greatly improved model fit, as indicated by Watanabe-Akaike information criterion, a commonly used Bayesian model comparison tool (7). The inference from the standard logistic regression analysis is likely too narrow and may result in type I errors. In fact, the standard approach suggests that three of the four effects (geographic distance, time between diagnosis dates, and same village) are statistically significant, whereas GenePair identifies only one (same village). When GenePair is fit to the “shuffled” data (the wrong person identifiers but the exact same data), the results nearly match the standard approach that ignores correlation, as expected. This suggests that GenePair does not artificially inflate uncertainty, and it can safely be used even with data for which the correlation between pairs is believed to be negligible. Overall, we recommend that future studies that leverage dyadic outcomes carefully consider the implications of ignoring correlation, as it can change the main conclusions of the analysis. Statistical methods for network data may be appropriate in this setting.

Figure 1. Posterior means and 95% credible intervals for three models: one that ignores correlation, as was done by Trevisi and colleagues (1) (light gray), one in which GenePair was used to address correlation in the data (black), and one in which GenePair was used, but data were shuffled to remove any correlation (medium gray).

Beyond our concern about the statistical analysis of the paired data presented by Trevisi and colleagues (1), we also question approaches for the identification of likely transmission pairs using only a SNP threshold as well as claims of directionality of transmission based solely on dates of participant diagnosis. In Trevisi and colleagues’ study, a binary threshold for labeling a pair as one that is connected through direct transmission was set at three SNPs or fewer, but in many M. tuberculosis genomic studies, multiple isolates are within transmission networks with little genetic diversity (8), thus leading to the potential designation of more than one individual as the source of another individual’s infection. Furthermore, given the variable timing of symptom onset and diagnosis relative to infection with M. tuberculosis, identifying the directionality of transmission on the basis of the first patient diagnosed in a pair may introduce misclassification. More robust methods for identifying the most likely transmission pairs, including using the full information in the genomic data and epidemiological data to inform inference about directionality of transmission, are available for these types of analyses (e.g., TransPhylo [9], outbreaker [10]). We note that even these methods leave considerable uncertainty about the identification of transmission pairs, and the value of such inference depends on completeness and approaches for sampling.

Author Contributions: T.C., C.C., and J.L.W. contributed to the conception and design. T.C. and J.L.W. wrote the first draft. J.L.W. performed simulations. T.C., CC., and J.L.W. edited and approved the final version.

Originally Published in Press as DOI: 10.1164/rccm.202404-0835RL on July 17, 2024

Author disclosures are available with the text of this letter at www.atsjournals.org.
==== Refs
References

1. Trevisi L Brooks MB Becerra MC Calderón RI Contreras CC Galea JT et al. Who transmits tuberculosis to whom: a cross-sectional analysis of a cohort study in Lima, Peru Am J Respir Crit Care Med 2024 210 222 233 38416532
2. Kenny DA Kashy DA Cook WL Dyadic data analysis New York Guilford Publications 2020
3. Warren JL Chitwood MH Sobkowiak B Colijn C Cohen T Spatial modeling of Mycobacterium tuberculosis transmission with dyadic genetic relatedness data Biometrics 2023 79 3650 3663 36745619
4. Hoff P Additive and multiplicative effects network models Statist Sci 2021 36 34 50
5. Hoff PD Random effect models for network data Breiger R Carley K Pattison P Dynamic social network modeling and analysis: workshop summary and papers Washington, DC National Academies Press 2003 303 312
6. Yang C Sobkowiak B Naidu V Codreanu A Ciobanu N Gunasekera KS et al. Phylogeography and transmission of M. tuberculosis in Moldova: a prospective genomic analysis PLoS Med 2022 19 e1003933 35192619
7. Watanabe S Opper M Asymptotic equivalence of Bayes cross validation and widely applicable information criterion in singular learning theory J Mach Learn Res 2010 11 3571 3594
8. Walker TM Ip CL Harrell RH Evans JT Kapatai G Dedicoat MJ et al. Whole-genome sequencing to delineate Mycobacterium tuberculosis outbreaks: a retrospective observational study Lancet Infect Dis 2013 13 137 146 23158499
9. Didelot X Fraser C Gardy J Colijn C Genomic infectious disease epidemiology in partially sampled and ongoing outbreaks Mol Biol Evol 2017 34 997 1007 28100788
10. Campbell F Didelot X Fitzjohn R Ferguson N Cori A Jombart T outbreaker2: a modular platform for outbreak reconstruction BMC Bioinformatics 2018 19 363 30343663
