
==== Front
Brief Bioinform
Brief Bioinform
bib
Briefings in Bioinformatics
1467-5463
1477-4054
Oxford University Press

10.1093/bib/bbae416
bbae416
Problem Solving Protocol
AcademicSubjects/SCI01060
Efficient test for deviation from Hardy–Weinberg equilibrium with known or ambiguous typing in highly polymorphic loci
https://orcid.org/0009-0005-5532-1578
Shkuri Or Department of Mathematics, Bar-Ilan University, Ramat Gan 5290002, Israel

https://orcid.org/0000-0003-1724-4979
Israeli Sapir Department of Mathematics, Bar-Ilan University, Ramat Gan 5290002, Israel

Tshuva Yuli Department of Mathematics, Bar-Ilan University, Ramat Gan 5290002, Israel

Maiers Martin CIBMTR (Center for International Blood and Marrow Transplant Research), and NMDP, Minneapolis, MN 55401-1206, USA

Louzoun Yoram Department of Mathematics, Bar-Ilan University, Ramat Gan 5290002, Israel

Corresponding authors: Or Shkuri, E-mail: orshkuri2000@gmail.com; Yoram Louzoun, E-mail: louzouy@math.biu.ac.il
9 2024
23 8 2024
23 8 2024
25 5 bbae41607 5 2024
01 7 2024
05 8 2024
© The Author(s) 2024. Published by Oxford University Press.
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited. For commercial re-use, please contact journals.permissions@oup.com

Abstract

The Hardy–Weinberg equilibrium (HWE) assumption is essential to many population genetics models. Multiple tests were developed to test its applicability in observed genotypes. Current methods are divided into exact tests applicable to small populations and a small number of alleles, and approximate goodness-of-fit tests. Existing tests cannot handle ambiguous typing in multi-allelic loci. We here present a novel exact test Unambiguous Multi Allelic Test (UMAT) not limited to the number of alleles and population size, based on a perturbative approach around the current observations. We show its accuracy in the detection of deviation from HWE. We then propose an additional model to handle ambiguous typing using either sampling into UMAT or a goodness-of-fit test test with a variance estimate taking ambiguity into account, named Asymptotic Statistical Test with Ambiguity (ASTA). We show the accuracy of ASTA and the possibility of detecting the source of deviation from HWE. We apply these tests to the HLA loci to reproduce multiple previously reported deviations from HWE, and a large number of new ones.

Hardy–Weinberg equilibrium
imputation algorithms
Gibbs sampling
Office of Naval Research 10.13039/100000006 N00014-23-1-2057 ISF 10.13039/100012579 870/20 DSI 10.13039/100016962
==== Body
pmcIntroduction

One of the most frequent assumptions regarding the allele pair compositions in a well-mixed population is the Hardy–Weinberg equilibrium (HWE). The HWE presumes random-pairing of alleles in the population. Formally, one assumes constant genetic variation in a given locus, and that the allele probability in the two chromosomes of each host is pairwise independent. The HWE has been tested in many settings [1], including among many others single nucleotide polymorphisms (SNPs) [2], multi-allele loci, homogeneous versus structured populations [3], polyploid [4] versus diploids. We here focus on a single polymorphic locus. A classical example of such a locus would be the human chromosome 6 HLA region [5]. The HWE assumption in HLA-based population genetics has been tested in hundreds of studies [6]. In this case, given an \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document} size sample from a population of size \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N_{P}$\end{document} with alleles taken from \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{1,...,N_{A}\}$\end{document} and unknown probabilities for the allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} - \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i)$\end{document}, and for the pair \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{i,j\}$\end{document} - \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document}, the HWE assumption is that

(1) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} \forall i,j \in \{1,...,N_{A}\}: p(i,j)=p(i) \cdot p(j). \end{eqnarray*}\end{document}

Note that one typically only computes the symmetric probabilities. In such a case, one would mark

(2) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} \forall i < j \in \{1,...,N_{A}\}: p(i,j)=2p(i) \cdot p(j);\:p(i,i)=p(i)^{2}. \end{eqnarray*}\end{document}

While \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document} is usually unknown, in most cases, one assumes that the observed allele pairs frequencies \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document} are a multinomial variable with the appropriate \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document}. This is often further approximated when \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document} is large by a binomial distribution for each \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document} using the appropriate \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document}. However, in some cases, one does not know for a given individual \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $k$\end{document} its precise allele pair. Instead, one may be given for the individual \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $k$\end{document} a distribution of possible candidate pairs, each with an estimated probability in this sample of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O_{k}(i,j)$\end{document}. One can then define the typing resolution score (TRS) [7] as follows:

(3) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} TRS(k)=\sum_{i,j} O_{k}(i,j)^{2} \end{eqnarray*}\end{document}

In the absence of ambiguity, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $TRS(k)=1$\end{document}, else \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $TRS(k)<1$\end{document}. Note that in this context, the distribution of

(4) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} O(i,j)=\sum_{k} O_{k}(i,j) \end{eqnarray*}\end{document}

is not multinomial anymore, and for a high enough ambiguity, and large enough \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document}, the central limit theorem holds for each \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i,j$\end{document}, and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document} is normally distributed, even if it is small.

Multiple methods were developed to test the deviation from HWE. Those can be divided into a few categories. The first is the large-sample size goodness-of-fit tests. These tests assume a large population and therefore assume normality of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document}. A very fundamental test in this group is the Chi-Square test [8]. Many variations of the Chi-Square were developed [9]. Such tests are only applicable to large populations.

An alternative is exact tests, based on a multinomial distribution assumption [10–12]. Such tests are more accurate than goodness-of-fit test tests for small populations but require extensive computing [13]. A newer exact test was introduced [14] which works more efficiently and is therefore applicable for larger populations, but it is only applicable to two alleles.

A third group is likelihood tests. One example can be found in Elston et al. [15] that works only for small samples. Another such test was proposed by Yu et al. [16] that assumes normality and therefore is not as accurate as the exact tests for small samples. Other alternatives are Bayesian tests [17, 18] and Monte-Carlo-based methods [13, 19].

Recently, Montoya-Delgado et al. [20] stated that current exact tests are not feasible for more than three alleles. They introduced a recursion-based exact test that avoids repeating calculations. However, even this test is computationally expensive for more than four alleles. To deal with more than four alleles, they used a permutation test for an approximate P-value. We here also include a permutation based test that is not limited by the number of alleles.

None of the tests above can account for ambiguity in the observations in multi-allelic loci, as will be further defined. Recent papers aim to solve the problem of ambiguity [21, 22]. However, all the proposed tests with ambiguity are limited to bi-allelic systems.

To summarize, goodness-of-fit tests require large populations, and exact tests are limited to small populations and a limited number of alleles. Beyond that, all existing methods cannot handle ambiguity in highly polymorphic loci. We propose two tests to solve for a single multi-allelic locus in diploids the two problems above. (A) The high cost of exact tests for un-ambiguous samples. (B) the variance estimate in goodness-of-fit tests in ambiguous samples.

To understand the effect of ambiguous typing, let us assume a set of two alleles: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $a$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $A$\end{document}, with equal probability. The homozygote probability \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $P(AA)=0.25$\end{document}, and in a sample of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document} alleles, the expected value of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(AA)$\end{document} is \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $0.25N$\end{document} and the variance is \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $0.25 \cdot 0.75 \cdot N$\end{document}. Now, assume instead that for each individual, we do not know the typing and assign \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O_{k}(AA)=0.25$\end{document}. In such a case, the expected value of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(AA)$\end{document} is still \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $0.25N$\end{document}. However, the variance is 0 (since all samples have \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O_{k}(AA)=0.25$\end{document}). One could propose to replace the unknown value, with a Bernoulli experiment, with the same probability, or produce a better estimate of the variance for the Wald test. We will show that this leads to different results.

Materials and methods

Allele pair simulation

Let us first assume that we observe the real allele pairs, (i.e. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{O(i,j)\}$\end{document} are known). We choose a value \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $0\leq \alpha \leq 1$\end{document} that represents the fit of the simulation to the HWE.

Produce a distribution \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{p(i)\}_{i}$\end{document} s.t. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\sum _{i=1}^{N_{A}} p(i)=1$\end{document} by generating a sequence of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $q_{i} \sim Uniform(0,2)$\end{document} for each allele and then normalizing the probability using a Softmax [23]: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $P=Softmax (Q)$\end{document}. To compute the allele pairs probabilities, we first compute the probabilities assuming HWE:

Calculate: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\forall{i,j} \in \{1,2,...,N_{A}\}$\end{document}: (5) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}& p_{0}(i,j)= \begin{cases} p(i)^{2}, & \text{if}\ i=j\\ 2 \cdot p(i)p(j), & \text{if}\ i<j\\ 0, & \text{if}\ i>j \end{cases}\end{align*}\end{document}

In the case of deviation from HWE, we add the deviation and renormalize: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\forall{i,j} \in \{1,2,,...,N_{A}\}$\end{document}: (6) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} p(i,j)= \begin{cases} (1-\alpha) \cdot p(i) + \alpha \cdot p_{0}(i, i), & \text{if}\ i=j\\ \alpha \cdot p_{0}(i, j) & \text{if}\ i<j\\ 0, & \text{if}\ i>j \end{cases} \end{eqnarray*}\end{document}

Finally,

(7) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} O(i,j)=Binomial(N_{0},p(i,j)) \end{eqnarray*}\end{document}

For the ambiguous case, we do not compute the binomial random variables probabilities. Instead, we compute a set of allele pairs with different probabilities for each sample. For each sample \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $k$\end{document} with real alleles \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{i,j\}$\end{document}, with probability \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho $\end{document}, we replace that by a group of all possible allele pairs, each with a probability \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{p(t,m)\}_{t \leq m}$\end{document}, as defined in Eq (6).

HLA genotype data

We studied 8078 224 donors from the National Marrow Donor Program (NMDP) registry [24]. The dataset provided for each donor an HLA typing and the 21 sub-populations that combine into five broad populations (S1 Table) if known, any other typed race was converted to unknown. Each donor was imputed to produce the 20 most likely 5-locus haplotype pair using MR-GRIMM [25]. We computed for each donor its TRS.

SNP data

We analyzed a dataset consisting of 13 258 SNPs located on chromosome 6 from the thousand genome project [26]. Each of these SNPs can be represented by one of only two possible nucleotides. Our dataset comprises genetic information from 986 individuals.

Code and server

The methods that were introduced in this paper can be accessed via our python package https://github.com/louzounlab/HWE_Tests, including a user-friendly website that runs this package hwetests.math.biu.ac.il. Moreover, the code for all the simulations can be found in https://github.com/louzounlab/HWE_Simulations.

Unambiguous Multi Allelic Test for certain alleles

The Unambiguous Multi Allelic Test (UMAT) test estimates the change in log-likelihood resulting from the random swap of two alleles toward achieving HWE. These swaps maintain the marginal distribution throughout the analysis, with deviations allowed for at most 4 alleles, each deviating by at most a single individual from the expected marginal distributions.

We evaluate how frequently the swapped distribution yields a lower probability than the current distribution, only after a certain amount since the first swaps are biased. In this study, we initiated swaps after 30 000 iterations and conducted a total of 100 000 iterations, each involving two swaps. These values were chosen for this analysis but can be adjusted as needed. The detailed algorithm is provided in Appendix A.

Asymptotic Statistical Test with Ambiguity

Assuming ambiguous typing, one can adopt two options. (A) Utilize a sampling method that follows the probabilities associated with each allele pair in each donor, integrating the UMAT test as described earlier. (B) Incorporate ambiguity into the goodness-of-fit test itself. For detailed implementation, refer to the algorithm outlined in Appendix B.

Models

Estimate of expected distribution

To estimate the expected distribution, we follow a single highly polymorphic locus with a large number of candidate alleles, or similarly, multiple-phased loci with a large number of haplotypes. For the sake of notation, we denoted both haplotype and alleles as alleles, since our analysis does not test for Linkage Disequilibrium (LD), but focuses solely on HWE. The observations, denoted as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O_{k}(i,j)$\end{document}, represent the probability that sample \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $k$\end{document} has alleles \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document}. We further define \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)=\sum _{k} O_{k}(i,j)$\end{document}.

In the absence of ambiguity, each sample has a single pair of alleles. In such a case, if the HWE assumption holds, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document} can be approximated by a binomial Random Variable (RV) with an expected value and variance of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N \cdot p(i,j)$\end{document}, where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document} is the sample size, for small enough \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document} (otherwise the variance is slightly smaller). This approximation is correct for large enough populations. On the other hand, if ambiguity is significant, for each pair \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $(i,j)$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document} becomes the sum of a large number of low probabilities. According to the central limit theorem, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document} will then approximately follow a normal distribution. We refer to the first scenario as the deterministic case (DC) and the second as the ambiguous case (AC).

In both scenarios, an estimate of the expected \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document} under HWE is required, despite observing only \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{O_{k}(i,j) \}_{i,j,k}$\end{document}. Assuming the sample size is sufficiently large to estimate marginal distributions with negligible error, we compute \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i)$\end{document} as

(8) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} p(i)=\frac{\sum_{k,j} O_{k}(i,j)}{N}, \end{eqnarray*}\end{document}

where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document}  denotes the total number of samples. This assumption holds even for small \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document}, as long as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\frac{1}{\sqrt{Np(i)}}<<1$\end{document}. We then utilize Eq 2 to calculate the expected \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document}, where here and all along, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document} represents the expected pair distribution given the observed marginal distributions \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i)$\end{document} and HWE. In both AC and DC cases, the expected value of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document} is \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N\cdot p(i,j)$\end{document}. However, as we will further show, the expected variance in the AC scenario is significantly lower than in the DC scenario.

Gibbs sampling based exact test for DC case—UMAT

In the DC case, one can compute the observed sample probability as

(9) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*} p(obs)&=\prod_{j \geq i} p(O_{i,j}|p_{i,j}) \end{align*}\end{document}

(10) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*} \forall{i \le j}:\:&p(O_{i,j}|p_{i,j})={N \choose O_{i,j}} \cdot p_{i,j}^{O_{i,j}} \cdot (1-p_{i,j})^{N-O_{i,j}}\end{align*}\end{document}

At this stage, instead of computing the entire distribution, we proposed to test the effect on the log-likelihood of performing a small perturbation, updating \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(k,l)$\end{document} to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(k,l)+1$\end{document}, assuming the change is small enough that its effect on the marginal distributions \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i)$\end{document} (and thus on the expected \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(i,j)$\end{document}) is negligible. The log ratio of the old to the new probabilities is

(11) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}& \begin{split} ln\left(\frac{p(O_{k,l}+1)}{p(O_{k,l})}\right) = &\ ln(N - O_{k,l})-ln(O_{k,l}+1)-ln(1-p_{k,l}) \\ & +ln(p_{k,l}) \end{split}\end{align*}\end{document}

Similarly, for the reduction of 1 from \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(k,l)$\end{document}

(12) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}& \begin{split} ln\left(\frac{p(O_{k,l}-1)}{p(O_{k,l})}\right) = &\ ln(O_{k,l})-ln(N-O_{k,l}+1)+ln(1-p_{k,l}) \\ & -ln(p_{k,l}) \end{split}\end{align*}\end{document}

Following an addition and a removal operation, the total sample size does not change. However, the marginal distributions may change. We thus proposed a sampling strategy that would ensure the marginal distribution is not changed along iterations of additional and removal. At each iteration, except for the first and last one, UMAT adds and removes a single individual, maintaining the marginal distribution, up to two samples along the entire analysis, and ensuring the HWE in the change (Fig. 1a). In short (see methods), we keep a pair of alleles where an individual was removed, with two initial values \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $m,n$\end{document} and a pair of alleles where an individual was added \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $k,l$\end{document}. We choose an allele from the second pair (say \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $l$\end{document}), and an additional allele chosen according to the distribution of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(t,l)\:\forall t$\end{document} (i.e. all the observed allele pairs with allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $l$\end{document}) and remove an individual and add it to the allele pair \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $m$\end{document} (or \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $n$\end{document}) and an allele chosen according to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(m)$\end{document}. The new individual was added following the HWE assumption, and there are still only two alleles with one missing individual and two alleles with one additional individual. We then compute the increase/decrease in the log-likelihood (Eq 11, 12) to produce the total change in the log-likelihood from the initial stage following each change. Note that at no stage do we compute the full probability distribution, only the changes. Finally, we check the fraction of changes where the cumulative change is larger than 0. If the HWE holds, this is a random walk around 0, and one can expect a distribution of the fraction of likelihoods above or below the initial values around \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $0.5$\end{document}. Otherwise, changes toward HWE should increase the likelihood. As such, most of the values will be above 0. We assume that a deviation of two samples from the marginal distribution has a negligible effect on the entire population, which holds for large enough populations.

Figure 1 DC case results; (a). visual representation of the swap of two alleles in the UMAT test; first sample an allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} according to the alleles distribution, then sample an allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document} according to the allele pairs distribution given allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} and subtract one individual with this pair (it must exist), then, sample a pair of alleles \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{k,l\}$\end{document} according to the HWE alleles distribution and add one individual to it; then, in each iteration, sample an allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $m$\end{document} according to the allele pairs distribution given allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $l$\end{document} and subtract one individual from the pair \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{m,l\}$\end{document}; sample an allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $t$\end{document} according to the allele marginal distribution and add one individual to the pair \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{t,j\}$\end{document}; this way, only the marginal observations of the alleles \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{t,k,m,i\}$\end{document} are changed; \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $t$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $k$\end{document} have one excessive marginal observation, as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $l$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $k$\end{document} had in the beginning and therefore take their place; note that we can randomly invert their order; similarly, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $m$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} are missing one marginal observation, as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} had in the beginning and therefore take their place in the next iteration; (b). 49 implementations of UMAT using the same observations in HWE, with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1.0$\end{document} (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $0 \leq \alpha \leq 1$\end{document} represents the fit of the simulation to the HWE, as described in Materials and methods), \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=1.e6$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $1000$\end{document} alleles; (c). 49 implementations of UMAT using the same observations with a slight deviation of HWE, with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =0.99$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=1.e6$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $1000$\end{document} alleles; (d). P-value results of UMAT for different allele numbers and alpha values, with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=1.e8$\end{document}; (e). P-value results of UMAT for different population sizes and alpha values, with 100 alleles; (f). elapsed time in seconds for running UMAT using different allele numbers and population sizes, with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1.0$\end{document}; as one can see, the population size has no effect on the run time; (g). scatter plot showing for 100 randomly chosen SNPs, the P-value results obtained with Chi-Square and UMAT.

Goodness-of-fit test for AC

We expanded our analysis to include ambiguous genetic typing scenarios, where the precise allele pair is unknown. For the sake of simplicity, we first described the deviation from the goodness-of-fit test when there is ambiguity in general, and only then discussed the HWE case. Consider \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document} samples are categorized into groups, each characterized by an unknown probability \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $q_{i}$\end{document}. We further assumed that given a “real” category \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document}, the observation is a distribution over all possible categories \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\eta _{i}(j) \forall j$\end{document}, where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\eta _{i}(j)$\end{document} is the probability of observing \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document} given the real category \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document}. The event of having \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $k$\end{document} cases of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} (although we may not observe them given the ambiguity) is a binomial random variable: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $ q_{i}^{k}=Binomial(q_{i},N,k)$\end{document} with expectation and variance (for small enough \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $q_{i}$\end{document}) of: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $E_{i}(k)=N \cdot q_{i}, V_{i}(k)=N \cdot q_{i}$\end{document} However, this is not the observed variable. We may observe \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document} instead of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document}. Assuming a set of samples, we can define a variable \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $z_{j}$\end{document}—the total mass of the observation \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document}. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $z_{j}$\end{document} is a random variable defined as a sum over donors and alleles:

(13) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} z_{j}=\sum_{l,i} I_{l}(i) \cdot \eta_{i}(j), \end{eqnarray*}\end{document}

where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $I_{l}(i)$\end{document} is the indicator of event \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} in donor \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $l$\end{document}. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} is the real allele, and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document} is the observed allele. We can compact the sum over donors to

(14) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} z_{j}=\sum_{i} q_{i}^{k} \cdot \eta_{i}(j) \end{eqnarray*}\end{document}

Here, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $z_{j}$\end{document} becomes a weighted sum of binomials, approximating a multinomial distribution with binomials. This approximation holds well for sufficiently large \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document}, despite the sum potentially differing from \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document}. The expected values of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $z_{j}$\end{document} are as before:

(15) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} E(z_{j})=\sum_{j} N \cdot q_{i} \cdot \eta_{i}(j). \end{eqnarray*}\end{document}

However, unlike the classical Chi-Square, the variances are smaller or equal to the mean:

(16) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} V(z_{j})=\sum_{j} N \cdot q_{i} \cdot \eta_{i}(j)^{2} \end{eqnarray*}\end{document}

Finally, we assumed that \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document} is large enough for the normal distribution approximation on \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document} (not on \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document}) to hold. In the deterministic case that \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\eta _{i}(j)=\delta _{i,j}$\end{document}, this converges to the classical Chi-Square test. Otherwise, per definition, the variance is smaller.

Allele pairs

Consider pairs of alleles \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $(l,k)$\end{document} and observed allele pairs \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $(m,n)$\end{document}. The index \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $i$\end{document} above is replaced by \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $(l,k)$\end{document}, and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document} is replaced by \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $(m,n)$\end{document}. In this case, we can define

(17) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*} &E(z_{m,n})=\sum_{l,k} E(l,k) \cdot p(m,n|l,k) \end{align*}\end{document}

(18) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*} &V(z_{m,n})=\sum_{l,k} E(l,k) \cdot p(m,n|l,k)^2\end{align*}\end{document}

However, the probabilities \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(m,n|l,k)$\end{document} are unknown. We thus use a simplifying assumption that \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(m,n|l,k) \sim p(m,n)$\end{document}, meaning the probability of observing \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $(m,n)$\end{document} is proportional to the unseen probability of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $(m,n)$\end{document} in the population. This is a strong assumption that is not always valid, and we have no clear evidence for its precision in different realistic scenarios. Under this assumption, we can estimate

(19) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*} &E(z_{m,n})= N \cdot p(m,n) \end{align*}\end{document}

(20) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*} &V(z_{m,n})=N \cdot p(m,n) \cdot \frac{\sum_{l,k} E(l,k) \cdot p(m,n|l,k)^{2}}{N \cdot p(m,n)}\end{align*}\end{document}

We can thus define a correction to the variance estimate as a sum over samples \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $t$\end{document}:

(21) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} \rho(m,n)=\frac{\sum_{t} p_{t}(m,n)^{2}}{\sum_{t} p_{t}(m,n)} \end{eqnarray*}\end{document}

and using the same approximation, we obtain

(22) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} V(m,n)=E(m,n) \cdot \rho(m.n) \end{eqnarray*}\end{document}

Finally, we revise the goodness-of-fit test to be

(23) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} Z=\sum_{m,n} \rho(m,n)^{-1} \cdot \frac{(0_{m,n}-E_{m,n})^{2}}{E_{m,n}} \end{eqnarray*}\end{document}

In the absence of ambiguity, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho (m,n)$\end{document} is simply 1. Note that the sum contains only non-negative values. If \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $m,n$\end{document} is never observed, then there is no correction, and arbitrarily, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho (m,n)=1$\end{document}.

Results

DC model validation

To test the accuracy of the DC model, we used simulated multi-allele samples and real-world SNP data (See methods). The simulation was based on the generation of allele pairs using the HWE assumption, and with a probability of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $1-\alpha $\end{document} choosing homozygotes instead of the current choice. We tested simulations with different population sizes between \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=1000$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=1.e8$\end{document}, with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $10-5000$\end{document} different alleles, and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =0.92, 0.99, 1.0$\end{document}. We first simulated a population of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=1.e6$\end{document} samples with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $1000$\end{document} different alleles in HWE (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1$\end{document}) and with a slight deviation \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =0.99$\end{document}. We performed 49 (7 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\times $\end{document} 7 plot) implementations of UMAT (Fig. 1b versus c). In HWE, the change in the log-likelihood is a random walk around 0 (B plot). In contrast, even a very small deviation of HWE (C plot) shows a consistent increase in the log-likelihood. To show that the results are not sensitive to the population size and allele number, we repeated the analysis with different allele numbers (Fig. 1d) and population size (Fig. 1e). In both cases, the fraction of time steps below the current likelihood is above 0.05 for \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =0$\end{document}, and lower than \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $1.e-3$\end{document} for \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =0.99$\end{document}. We repeated the analysis for both smaller populations and different allele numbers (S1 Fig A, B for \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1.0, 0.92$\end{document}, respectively).

We tested the run time of UMAT. This run time is only based on the need to choose the alleles, which is linear in the number of alleles, and independent of the population size (Fig. 1f).

We further tested the model on bi-allelic SNP data from the 1K genome project. The population may deviate from HWE. We analyzed 100 SNPs and compared the results of UMAT and a classical Chi-Square (given the large population size and the low number of alleles). One can see an excellent fit between the P-value of UMAT and of the Chi-Square (Fig. 1g). However, in general, the Chi-square gives a slight overestimate of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p$\end{document}.

Validation of AC method

Two approaches can be proposed to test the HWE assumption. First, one can sample the distribution of allele pairs in each individual to produce a single pair according to the probability of each candidate pair. Assume for example that a given sample has two candidate pairs with probabilities of 0.3 and 0.7, respectively. One can choose the first with a probability of 0.3 or the second with a probability of 0.7. One can then apply UMAT. To test the sampling, we used the simulation presented above in the DC case and added ambiguity by assigning a probability per sample to replace the real pair with a set of probabilities for possible pairs (See methods). Again, the test is highly accurate with practically no FP and a very accurate detection even for a very small deviation from HWE (Fig. 2a).

Figure 2 AC case results; (a). P-value results of UMAT with ambiguity for different \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho $\end{document} values (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho $\end{document} represents the probability for a sample to have ambiguous typing, as described in Materials and methods) and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha $\end{document} values, with 100 alleles and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=10,000$\end{document}; (b)–(d). scatter plot of real versus estimated variance for simulated data with 50 alleles, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=1.e4, \alpha =1, \rho =0,0.2,0.4$\end{document} (respectively); each dot is an allele pair \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{i,j\}$\end{document}; the x-axis represents the sampled variance over \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\{O_{k}(i,j)\}_{k}$\end{document}, the y-axis represents either the value of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $O(i,j)$\end{document} or the corrected denominator used in ASTA: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $E(i,j) \cdot \rho (i,j)$\end{document}; (e). fraction of positive results for \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =0.92, 0.96, 1.0$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho =0.0, 0.1, 0.2, 0.4$\end{document} and each test: raditional Chi-Squared, ASTA, and Chi-Squared with sampling (i.e. for each person sampling certain alleles given all his possible allele pair observations and then using a traditional Chi-Squared); the results are out of 300 simulations with five alleles and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=2000$\end{document}; (f). we simulated data (see Methods) using \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $100$\end{document} alleles, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N=10\,000, \alpha =1.0$\end{document}; here all the data are in HWE, except from the pairs containing the \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $0$\end{document} allele (top bar); each bar corresponds to a specific allele and shows the normalized statistic (the sum of Chi-Square term with correction only on pairs containing this allele, and divided by the Degree of Freedom (DOF): the number of pairs containing this allele minus 1), as well as the \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $-log(p{\_ }value)$\end{document} for this allele, here the \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p{\_ }value$\end{document} is calculated using the statistic and DOF of the allele.

An alternative method is to improve the variance estimate in the goodness-of-fit test test. To test the deviation of the variance from the binomial estimate, we compared the variance and the estimated value for each allele pair \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $(m,n)$\end{document} as a function of the ambiguity (Fig. 2b-d). As expected, when \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1$\end{document}, the corrected variance estimate is equivalent to the observed value and is approximately the expected value. However, when the ambiguity is increased, the variance is much lower than the expected value but is similar to the corrected variance estimate.

We further tested the regular Chi-square and Asymptotic Statistical Test with Ambiguity (ASTA) for \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =0.92-1.0$\end{document} and four different values of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho \; (0.0, 0.1, 0.2, 0.4)$\end{document}, we ran the simulation 300 times and averaged the fraction of significance results (each result is either 1 for significance or 0) at the \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $P=0.05$\end{document} level. We used a population of size \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N_{0}=2000$\end{document} and five alleles (Fig. 2e).

Both the old Chi-square and ASTA coincide for \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho =0.0$\end{document} since the term of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $V(i,j)$\end{document} is the same in the denominator. Also, as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha $\end{document} approaches 1.0 (i.e. original observations get closer to HWE), and as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho $\end{document} gets higher, the classical Chi-Square reports a deviation from HWE, while the real observations are in HWE. ASTA does not produce more false positives than randomly expected. Moreover, it detects much more precisely the deviation from HWE, even for small deviations as was the case for UMAT. Results for a higher number of alleles and larger populations are even more accurate and are given in the Appendix (S2 Fig).

A clear advantage of this model is the possibility to detect the alleles associated with the deviation from HWE. Instead of summing over all allele pairs, one can sum the Chi-Square term only on pairs containing allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $t$\end{document}.

To test this, we performed simulations where we deliberately selected an allele, referred to as allele 0, to deviate from HWE. from HWE, such that only allele pairs containing the allele called 0 are not in HWE. For pairs containing this allele, we choose the HWE distribution with probability \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha $\end{document}, and the homozygotes otherwise. To test that, we set \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =0.7$\end{document} and perform for every allele \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $j$\end{document}

(24) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{eqnarray*} p(0,j)= \begin{cases} (2\alpha - 1) \cdot p(0)^{2} + 2(1-\alpha) \cdot p(0), & \text{if}\ j=0\\ \alpha \cdot p_{0}(i, j) & \text{if}\ j>0 \end{cases}\end{eqnarray*}\end{document}

We then tested whether the deviation from HWE was detected and whether the appropriate allele was detected (Fig. 2f). One can see that the main deviation is for the chosen allele (top row). The same holds for simulations with different numbers of alleles and population sizes.

Figure 3 Deviation of alleles from HWE; each column is one of the five broad US populations; each row is a locus (A, B, C, DQB1, DRB1); within each subplot the bar is the ASTA score divided by the DOF; the color represents the log P value; deeper colors are more significant; the bars are ordered by decreasing significance; note that the scale of each plot is very different, but the coloring of the bar is consistent among all subplots.

Deviation of HWE in HLA

To test the deviation from HWE in a large-scale human population, with a large number of alleles, we analyzed the HLA allele frequency distribution from 8.1 million US donors [27]. The HLA locus has five classical loci (A,B,C,DRB1,DQB1), and the population is divided into five broad populations (CAU—White/European, HIS-Hispanic or Latino, AFA-Black/African American, API—Asian or Pacific Islander and NAM—Native American). Each broad population is further divided into sub-populations. We computed for each broad population and each locus the alleles that deviate the most (highest ASTA score) from HWE (Fig. 3). This measurement reproduces multiple previous results. For example, in the CAU populations, the alleles that deviate the most are related to an Ashkenazi Jew haplotyope -A*26:01 B*38:01 C*12:03 DQB1*03:02 DRB1*04:02 [28]. Another example would be the A*34:01 that has a clear north-south separation in the Asian population [29]. The top A allele (A*01:03) and top B allele (B*41:01) in AFA are both from eastern and northern Africa and not sub-Saharan African. DRB1*04:07 the most out of HWE DRB1 in the AFA population. However, this is a Colombian allele that has been mixed with the AFA population, mainly in Colombia [27, 30].

To further study the total deviation from HWE in the total population, we compared all three models (ASTA, Chi-Square, Chi-Square with sampling) on all broad and sub-groups and for all loci (Fig. 4). The population with the strongest deviation from HWE is the Asian Pacific population (API), since it includes a combination of multiple sub-populations, including Chinese and Indian. Within it the API population the Indian (AINDI) and mixed Asian (SCSEAI) have the highest deviation in all loci. The homogeneous populations, like Japan, Korea and Vietnam have consistently very low deviation from HWE. Interestingly, the Mexican population (MSWHIS) deviates the most in the A locus.

Finally, to demonstrate the improvement of ASTA over the Chi-Squared test, we simulated various datasets that remain in Hardy-Weinberg equilibrium (HWE) but with different levels of uncertainty. The simulations were generated as described in (Fig. 5), where we calculated the Chi-squared, ASTA, and inverse CDF of Chi-squared statistics using a 0.5 significance level. For each simulation, we varied the degrees of freedom (DOF) by starting with the first two allele pairs (1 DOF) and incrementally including additional pairs, recalculating the statistics as the DOF increased, until all allele pairs were included. As shown in the results, when uncertainty is low, the Chi-Squared statistic closely follows the inverse CDF of the Chi-Squared distribution. However, as uncertainty increases, the Chi-Squared statistic tends to decrease relative to the inverse CDF. In contrast, ASTA consistently aligns with the inverse CDF, except in scenarios with very large degrees of freedom.

Figure 4 Each subplot is a different locus (A,B,C,DQB1, DRB1); the different colors represent the log base 10 of: the scores divided by the DOF; the ASTA score (orange) is slightly higher than the classical Chi-Square (blue) and from the sampling (green); note that in this case, the ambiguity is limited; as such, the differences are not very large; the top populations are the sub-populations, followed by the broad populations, followed by the entire donor registry; each detailed population is colored according to the broad population.

Figure 5 Statistics of Chi-Squared, ASTA, and the inverse CDF of Chi-Squared using a 0.5 significance level; each statistic is calculated for different degrees of freedom (DOF) values, ranging from 1 to the number of allele pairs minus 1; the statistics are first calculated for the first two allele pairs (1 DOF), and subsequently, one additional pair is included to recalculate the statistics (increasing the DOF by 1) until all pairs are included; observations for each plot are generated with the following parameters: (a). \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $50$\end{document} alleles, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $50\,000$\end{document} population size, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1.0$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho =0.1$\end{document}; (b). \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $100$\end{document} alleles, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $200\,000$\end{document} population size, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1.0$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho =0.2$\end{document}; (c). \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $150$\end{document} alleles, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $250\,000$\end{document} population size, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1.0$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho =0.15$\end{document}; (d). \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $200$\end{document} alleles, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $400\,000$\end{document} population size, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\alpha =1.0$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\rho =0.3$\end{document}.

Conclusion

The HWE assumption is crucial for many population genetics models including, among many others, Expectation Maximization (EM) based haplotype frequency estimate [31, 32], imputation, admixture algorithms [25, 33, 34], and any model containing an estimate of probability of allele pairs. We have here focused on the specific context of multi-allelic loci in diploids. We propose in this context two algorithms for the estimate of deviation from HWE. The first (UMAT) is a method to test the deviation of the exact distribution from the one expected in HWE, and the second (ASTA) is a goodness-of-fit test that can handle ambiguous typing. UMAT is based on a perturbative approach, where we test the effect of randomly swapping pairs of alleles on the likelihood. ASTA is based on the correction of the Chi-Square test to reflect the variance with ambiguous typing. We show the accuracy of both models using extensive simulation and real-world cases. ASTA is per definition an approximation, but it has the advantage of detecting the alleles/haplotypes most associated with the deviation from HWE.

We have tested both methods using the well-known HLA locus in the human chromosome 6, and shown a significant deviation from HWE in practically all genes and all sub-populations, with a larger deviation, in broad populations than in sub-populations. We further detected the alleles most associated with deviation from HWE, and have shown it reproduces multiple previous reports on deviations from HWE in HLA.

The main advantage of UMAT and ASTA over existing methods is the possibility of estimating the deviation from HWE from thousands or tens of thousands of alleles/haplotypes, and with no limit on the population size. Both methods have limitations that are handled by other models. Those include among others the limitation to a single locus and the lack of a solution for a structured population [3]. In the case of ASTA, there are other limitations, including the normal approximation, which is only valid if we sum enough elements. Note that for ambiguous typing, the total number of elements in the sum can be much larger than the expected value. To handle that, we allow the user to limit the test to a minimal denominator, and ignore allele/haplotype pairs with too low denominators. Another quite stringent assumption is that the expectation of the observed value is indeed proportional to the allele pair probability. This is trivial in the regular goodness-of-fit test. However, in the ambiguous case, the ambiguity may be biased toward specific pairs. We currently have no solution if this assumption is not at least approximately valid.

Key Points

Hardy–Weinberg equilibrium estimate is crucial for many bioinformatics tools. However, there is no accurate estimator for the HWE in large multi-locus populations.

We present two algorithms to solve that. UMAT based on GIBBS sampling, and a corrected goodness-of-fit test algorithm named ASTA.

UMAT can detect very small deviations from HWE in a short time even with a very large number of alleles and in large populations.

ASTA can further detect the alleles associated with deviation from HWE.

We applied those to an 8 million sample of HLA genotypes, and detected novel alleles associated with deviation from HWE.

Supplementary Material

S1_fig_bbae416

S2_fig_bbae416

S3_fig_bbae416

S4_fig_bbae416

S5_fig_bbae416

S6_fig_bbae416

S7_fig_bbae416

S1_table_bbae416

Acknowledgments

The authors thank the anonymous reviewers for their valuable suggestions. We thank Miriam Beller for the English editing.

Funding

The work of M.M. and Y.L. was partially funded by the Office of Naval Research grant (N00014-23-1-2057) (https://govtribe.com/vendors/national-marrow-donor-program1msy6). The work of S.I. was funded by ISF grant 870/20 (https://www.isf.org.il/#/) and by a DSI Grant (https://dsi.biu.ac.il/).

 

Competing interests: No competing interest is declared.

Data availability

The code used for simulations is available at https://github.com/louzounlab/HWE_Simulations. The SNP data are available at KAGGLE: https://www.kaggle.com/datasets/extraflash/snipping-data. The HLA frequencies data are availabale at https://www.allelefrequencies.net/. The HLA genotypes are available upon request from Martin Maiers at the NMDP.

Appendix

A. UMAT

B. ASTA
==== Refs
References

1 Hou  Y, Prinz  M, Staak  M. Comparison of different tests for deviation from hardy-weinberg equilibrium of ampflp population data. In: Bär W, Fiori A, Rossi U (eds.), Advances in Forensic Haemogenetics: 15th Congress of the International Society for Forensic Haemogenetics (Internationale Gesellschaft für forensische Hämogenetik eV), Venezia, 13–15 October 1993, pp. 511–4. Springer-Verlag, Berlin Heidelberg, 1994. 10.1007/978-3-642-78782-9_141.
2 Rohlfs  RV, Weir  BS. Distributions of hardy–weinberg equilibrium test statistics. Genetics  2008;180 :1609–16. 10.1534/genetics.108.088005.18791257
3 Hao  W, Storey  JD. Extending tests of hardy–weinberg equilibrium to structured populations. Genetics  2019;213 :759–70. 10.1534/genetics.119.302370.31537622
4 Sun  L, Gan  J, Jiang  L. et al.  Recursive test of hardy-weinberg equilibrium in tetraploids. Trends Genet  2021;37 :504–13. 10.1016/j.tig.2020.11.006.33341281
5 Breuning  MH, van den Berg-Loonen  EM, Bernini  LF. et al.  Localization of hla on the short arm of chromosome 6. Hum Genet  1977;37 :131–9. 10.1007/BF00393575.885534
6 Sakaue  S, Gurajala  S, Curtis  M. et al.  Tutorial: a statistical genetics guide to identifying hla alleles driving complex disease. Nat Protoc  2023;18 :2625–41. 10.1038/s41596-023-00853-4.37495751
7 Paunić  V, Gragert  L, Schneider  J. et al.  Charting improvements in us registry hla typing ambiguity using a typing resolution score. Hum Immunol  2016;77 :542–9. 10.1016/j.humimm.2016.05.002.27163154
8 Karl Pearson  X . On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. London Edinburgh Philos Mag & J Sci  1900;50 :157–75.
9 Emigh  TH . A comparison of tests for hardy-weinberg equilibrium. Biometrics  1980;36 :627–42. 10.2307/2556115.25856832
10 Levene  H . On a matching problem arising in genetics. Ann Math Stat  1949;20 :91–4. jstor. 10.1214/aoms/1177730093. http://www.jstor.org/stable/2236806  (21 August 2023, date last accessed).
11 Chapco  W . An exact test of the hardy-weinberg law. Biometrics  1976;32 :183–9  jstor. 10.2307/2529348  (21 August 2023, date last accessed).1276367
12 Haldane  j . An exact test for randomness of mating. J Genet  1954;52 :631–5. 10.1007/BF02981502.
13 Guo  SW, Thompson  EA. Performing the exact test of hardy-weinberg proportion for multiple alleles. Biometrics  1992;48 :361–72. 10.2307/2532296.1637966
14 Wigginton  JE, Cutler  DJ, Abecasis  GR. A note on exact tests of hardy-weinberg equilibrium. Am J Hum Genet  2005;76 :887–93. 10.1086/429864.15789306
15 Elston  RC, Forthofer  R. Testing for hardy-weinberg equilibrium in small samples. Biometrics  1977;33 :536–42. 10.2307/2529370.
16 Chang  Y, Zhang  S, Zhou  C. et al.  A likelihood ratio test of population hardy-weinberg equilibrium for case-control studies. Genet Epidemiol Society  2009;33 :275–80.
17 Lindley  D . Statistical inference concerning hardy-weinberg equilibrium. Bayesian Stat  1988;3 :307–26.
18 Montoya-Delgado  LE, Irony  TZ, Pereira  CA d B. et al.  An unconditional exact test for the hardy-weinberg equilibrium law: sample-space ordering using the bayes factor. Genetics  2001;158 :875–83. 10.1093/genetics/158.2.875.11404348
19 Lazzeroni  LC, Lange  K. Markov chains for Monte Carlo tests of genetic equilibrium in multidimensional contingency tables. Ann Stat  1997;25 :138–68. 10.1214/aos/1034276624.
20 Graffelman  J, Ortoleva  L. A network algorithm for the x chromosomal exact test for hardy–weinberg equilibrium with multiple alleles. Mol Ecol Resour  2021;21 :1547–57. 10.1111/1755-0998.13373.33687797
21 Shriner  D . Approximate and exact tests of hardy-weinberg equilibrium using uncertain genotypes. Genet Epidemiol  2011;35 :632–7. 10.1002/gepi.20612.21922537
22 Kwong  AM, Blackwell  TW, LeFaive  J. et al.  Robust, flexible, and scalable tests for hardy–weinberg equilibrium across diverse ancestries. Genetics  2021;218 :iyab044. 10.1093/genetics/iyab044.33720349
23 Bridle  J . Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters. In: Touretzky D (ed.), Adv Neural Inf Process Syst. Morgan-Kaufmann, Great Malvern, Worcestershire, UK, 1989;2 .
24 Gragert  L, Spellman  SR, Shaw  BE. et al.  Unrelated stem cell donor hla match likelihood in the us registry incorporating hla-dpb1 permissive mismatching. Transplantation and cellular. Therapy  2023;29 :244–52. 10.1016/j.jtct.2022.12.027.
25 Israeli  S, Gragert  L, Madbouly  A. et al.  Combined imputation of hla genotype and self-identified race leads to better donor-recipient matching. Hum Immunol  2023;84 :110721. 10.1016/j.humimm.2023.110721.37867095
26 Fairley  S, Lowy-Gallego  E, Perry  E. et al.  The international genome sample resource (igsr) collection of open human genomic variation resources. Nucleic Acids Res  2020;48 :D941–7. 10.1093/nar/gkz836.31584097
27 Gragert  L, Madbouly  A, Freeman  J. et al.  Six-locus high resolution hla haplotype frequencies derived from mixed-resolution dna typing for the entire us donor registry. Hum Immunol  2013;74 :1313–20. 10.1016/j.humimm.2013.06.025.23806270
28 Klitz  W, Loren Gragert  M, Maiers  MF-V. et al.  Genetic differentiation of jewish populations. Tissue Antigens  2010;76 :442–58. 10.1111/j.1399-0039.2010.01549.x.20860586
29 Mack  St, Alicia  S-M, Meyer  D. et al.  13th international histocompatibility workshop anthropology/human genetic diversity joint report. In: Meyer D (ed.), Immunobiology of the Human MHC: Proceedings of the 13th International Histocompatibility Workshop and Conference, Seattle: Int. Histocompatibility Working Group Press, 2006, 1:564–79, 01.
30 Trachtenberg  EA, Keyeux  G, Bernal  JE. et al.  Results of expedition humana: I. Analysis of hla class ii (drb1–dqa1-dqb1-dpb1) alleles and dr-dq haplotypes in nine amerindian populations from Colombia. Tissue Antigens  1996;48 :174–81. 10.1111/j.1399-0039.1996.tb02625.x.8896175
31 Kuk  AYC, Zhang  H, Yang  Y. Computationally feasible estimation of haplotype frequencies from pooled dna with and without hardy–weinberg equilibrium. Bioinformatics  2009;25 :379–86. 10.1093/bioinformatics/btn623.19050036
32 Israeli  S, Gragert  L, Maiers  M. et al.  Hla haplotype frequency estimation for heterogeneous populations using a graph-based imputation algorithm. Hum Immunol  2021;82 :746–57. 10.1016/j.humimm.2021.07.001.34325951
33 Deng  H-W, Chen  W-M, Recker  RR. Population admixture: detection by hardy-weinberg test and its quantitative effects on linkage-disequilibrium methods for localizing genes underlying complex traits. Genetics  2001;157 :885–97. 10.1093/genetics/157.2.885.11157005
34 Pearman  WS, Urban  L, Alexander  A. Commonly used hardy–weinberg equilibrium filtering schemes impact population structure inferences using radseq data. Mol Ecol Resour  2022;22 :2599–613. 10.1111/1755-0998.13646.35593534
