
==== Front
Commun Biol
Commun Biol
Communications Biology
2399-3642
Nature Publishing Group UK London

39232115
6769
10.1038/s42003-024-06769-3
Article
Extracting regulatory active chromatin footprint from cell-free DNA
Lai Kevin
Dilger Katharine
Cunningham Rachael
Lam Kathy T.
Boquiren Rhea
Truong Khiet
Louie Maggie C.
Rava Richard
http://orcid.org/0009-0000-4159-0084
Abdueva Diana diana.abdueva@aqtual.com

AQTUAL Inc., 31145 San Antonio Street, Hayward, CA 94544 USA
4 9 2024
4 9 2024
2024
7 108628 2 2024
21 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Cell-free DNA (cfDNA) has emerged as a pivotal player in precision medicine, revolutionizing the diagnostic and therapeutic landscape. While its clinical applications have significantly increased in recent years, current cfDNA assays have limited ability to identify the active transcriptional programs that govern complex disease phenotypes and capture the heterogeneity of the disease. To address these limitations, we have developed a non-invasive platform to enrich and examine the active chromatin fragments (cfDNAac) in peripheral blood. The deconvolution of the cfDNAac signal from traditional nucleosomal chromatin fragments (cfDNAnuc) yields a catalog of features linking these circulating chromatin signals in blood to specific regulatory elements across the genome, including enhancers, promoters, and highly transcribed genes, mirroring the epigenetic data from the ENCODE project. Notably, these cfDNAac counts correlate strongly with RNA polymerase II activity and exhibit distinct expression patterns for known circadian genes. Additionally, cfDNAac signals across gene bodies and promoters show strong correlations with whole blood gene expression levels defined by GTEx. This study illustrates the utility of cfDNAac analysis for investigating epigenomics and gene expression, underscoring its potential for a wide range of clinical applications in precision medicine.

Enrichment and analysis of regulatory-active chromatin cfDNA fragments in blood offer a non-invasive method for monitoring dynamic changes in epigenetic and gene expression patterns, underscoring its potential utility in precision medicine.

Subject terms

Epigenomics
Transcriptomics
Personalized medicine
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

The discovery of circulating DNA and RNA in the plasma of healthy individuals and patients was made by Mandel and Metais in 19481. In 1966, Tan et al. observed an anomalous pattern of cell-free DNA (cfDNA) in patients who suffered from systemic lupus erythematosus2, demonstrating the potential utility of cfDNA for detecting disease. Additional early discoveries by Fournier3 and Stroun et al.4 of cfDNA in transplant and cancer patients, respectively, further suggested the potential utility of cfDNA as a useful non-invasive biomarker for diseases. Unlike DNA found within intact cells, cfDNA is released into the circulation when cells die either as a result of routine cell turnover or triggered by a specific disease5,6. The size distribution of cfDNA is not uniform, with reported fragment lengths of less than 100 bp to over 5000bp7. This size heterogeneity is due to the different mechanisms in which DNA is released from cells, the protection of the DNA by regulatory proteins, and subsequent degradation both in the blood and post-blood draw8.

Nucleosomes play a crucial role in the structure and packaging of DNA, including cfDNA. As the fundamental units of chromatin9, nucleosomes— together with their epigenetic modifications— regulate transcription by modulating the accessibility of regulatory factors to DNA. Historically, chromatin modifications were studied using chromatin immunoprecipitation-sequencing (ChIP-seq)10 and the assay for transposase-accessible chromatin using sequencing (ATAC-seq)11. These methods, while informative, are cumbersome due to their complex protocols and extensive sample preparation requirements. Consequently, their complexity and resource-intensive needs render them less suitable for routine clinical use.

The presence of nucleosomes in cfDNA has important implications for its stability, accessibility, and potential applications. Additionally, bound regulatory proteins (e.g. transcription factors) can protect the cfDNA from degradation, giving rise to cell-free chromatin that retains both the nucleosome and regulatory protein structure of the parent cells. Cell-free active chromatin (cfDNAac) contains valuable information regarding the tissue of origin, and molecular and cellular states. The added layer of the regulatory protein interactions associated with cfDNAac offers deeper biological insights compared to the cfDNA associated with only nucleosomes (cfDNAnuc), suggesting an enhanced role for cfDNAac in precision medicine and molecular diagnostics. This underscores the significance of isolating and examining this distinctive cfDNA population to fully exploit its comprehensive biomarker potential.

The extraction of the cfDNA from plasma can be accomplished using established off-the-shelf kits, but these methods often enrich for the shorter cfDNAnuc fragments associated with a mono-nucleosome, potentially losing information present in other cfDNA fragments12. A further complication is that the size of the cfDNA is often altered pre-analytically by both the experimental collection and initial processing techniques. Specifically, collecting cfDNA in different blood collection tubes, with or without chemical protectants, potentially determines the observed size distribution of the cfDNA fragments. Therefore, controlling pre-analytical steps is critical to extracting the complete spectrum of cfDNA fragments, including those found in active chromatin.

In this study, we have improved and optimized the cfDNA extraction process to more effectively capture the full spectrum of cfDNA fragment sizes present in blood. In addition, we developed computational techniques for analyzing all available cfDNA fragments and have demonstrated that we can significantly enrich cfDNAac signals with our improved methods. These active chromatin fragments provide information on gene expression and regulation, including the identification of functional elements within the cfDNAac. Our study illustrates that cfDNAac counts correlate with RNA polymerase II pausing profiles; and cfDNAac is effective in detecting diurnal expression patterns of circadian genes, demonstrating its capability in monitoring dynamic physiological changes in gene expression. Finally, the results of this study demonstrate the ability of cell-free chromatin analysis to recapitulate ChIP-derived data from the Encyclopedia of DNA Elements (ENCODE) project13 and highlight the promise of this approach for multiple clinical applications.

Results

The chromatin capture workflow enriches cfDNA fragments in regulatory regions

Blood was collected from twenty healthy individuals (see Methods) for this study. Cell free DNA was extracted from 2-mL of plasma from a single individual (labeled AS-003) across 14 time-points using an off-the-shelf kit following the manufacturer’s instructions. Additionally, 2-mL of plasma from five individuals (AS-001, AS-002, AS-003, AS-004, and AS-005) across the 14 timepoints were processed using our modified chromatin capture workflow (Supplementary Table S1). In both workflows, magnetic nanoparticles were used to bind to the cfDNA, and the lysis and binding conditions were optimized to enrich for active chromatin fragments.

Our results show that while off-the-shelf cfDNA extraction methods primarily enrich for shorter cfDNAnuc fragments encompassing the mono- and di-nucleosome cfDNA profile, our chromatin capture workflow is optimized to enrich for the longer cfDNAac fragments associated with active regulatory proteins (Fig. 1, Supplementary Data 1). The median fragment length distribution from the off-the-shelf workflow has a dominant peak at ~167 bp and a smaller second peak at ~334 bp, representing the mono- and di-nucleosome profile, respectively9,14 (Fig. 1b). Approximately 77% of fragments from this workflow are less than 200 bp. In contrast, our chromatin capture workflow produces a shifted fragment length distribution where the first peak is approximately 10 bp longer, and only eight percent of the fragments are less than 200 bp (Fig. 1b). As a result, the fragment size distribution in our workflow is considerably enriched for longer fragments (Fig. 1b). In addition, the fragment size distribution associated with mono-, di- and tri-nucleosome profiles is also wider, suggesting that the cfDNA fragments are not protected solely by nucleosomes (i.e. cfDNAnuc) but also by other protein-bound regulatory complexes (i.e. cfDNAac).Fig. 1 Active chromatin capture workflow enriches for regulatory-active chromatin fragments in plasma.

a Blood was collected, and plasma separation completed within 24 hours. cfDNA extraction using the off-the-shelf workflow followed the Apostle MiniMax cfDNA protocol to enrich for nucleosomal fragments (cfDNAnuc). The active chromatin capture workflow used optimized extraction and library preparation conditions to enrich for regulatory-active chromatin fragments (cfDNAac), as described in the Methods section. Prepared libraries were then sequenced on a 200 cycle NovaSeq 6000 S4 flowcell. b Size distributions of sequenced cfDNA fragments from subject AS-003, extracted using the off-the-shelf workflow (gray) and the chromatin capture workflow (red).

The deconvolution of cfDNA signal derived from active chromatin was accomplished by unsupervised clustering of fragment count profiles for different fragment size groups or bins (see Methods). After deconvolution, approximately 52% of fragments were defined as cfDNAac (Supplementary Table S2). In the following sections, we examine the value of these fragments for measuring regulatory elements across the genome.

cfDNAac fragments correlate with ENCODE mappings

ChIP-seq and its variations have played an enormous role in mapping the transcriptome and its associated regulatory elements in the ENCODE project13. Recently, the ENCODE consortium has expanded their encyclopedia with new experimental datasets that include the Roadmap Epigenomics database15 to better understand the functional elements in both human and mouse genomes16. With the hypothesis that the cfDNA extracted by our chromatin capture workflow is enriched for regulatory protein-protected fragments, we examined whether these cfDNAac fragments could recapitulate regulatory elements mapped in ENCODE.

In order to test this premise, we compared the cfDNAac fragment distribution with ENCODE mappings13 at active promoters (Fig. 2, Supplementary Data 1). The cfDNAac fragment distribution clearly overlaps with ENCODE mappings for several active promoter regions (Fig. 2a). In contrast, the cfDNAnuc fragments are not enriched in these regions. These results suggest that the cfDNAac fragments are derived from transcription factor (TF)-bound complexes within the genome’s regulatory regions and are protected from rapid degradation. The cfDNAac fragments also show an enrichment of signal around the transcription start sites (TSS)17 identified by ChIP-seq of histone modifications known to identify promoter regions18 (Fig. 2b).Fig. 2 Epigenetic profiling of regulatory-active chromatin fragments in plasma.

a Genome browser view of regulatory-active chromatin signal on a segment of chromosome 12 representing a promoter region. The first track represents the regulatory region protected from DNase degradation. The subsequent 5 tracks are GM12878 ChIP-Seq –log(p-value) signals for selected histone modifications (H3K4me1, H3K4me2, H3K4me3, H3K27ac, and H3K36me3). The last two tracks represent enrichment of nucleosomal (cfDNAnuc) and regulatory-active chromatin (cfDNAac) genome-wide signals. Signals were calculated by log2-fold change between the two fragment classes. b Comparison of the average signal distribution surrounding TSS from EPDnew with a +/−10 kb flank. The black arrow shows the gene orientation. The scale of ChIP-Seq distributions is –log(p-value), while the cfDNAnuc and cfDNAac distributions are scaled coverage (see Methods). c Pearson correlation between the median number of cfDNAnuc (gray) and cfDNAac fragments (red) to density of H3K4me1 narrow peaks across the genome in 1 Mb segments from all sequenced plasma samples (n = 70).

Extending the analysis across the genome, cfDNAac fragments are enriched in CpG islands and CpG shores compared to cfDNAnuc fragments (Supplementary Fig. S1). These data also indicate that no enrichment is observed in intronic and repressed regions across the genome for cfDNAac. To investigate whether we only enrich fragments with higher percent GC, we first stratified FANTOM5 TSS locations across the full range of percent GC (0–20th percentile, 20th–40th percentile, 40th–60th percentile, 60th–80th percentile, and 80th–100th percentile) and compared the raw fragment counts at these locations (Supplementary Fig. S2). The GC distribution is not uniform across promoters, with the interquartile falling between 56–69%. To minimize bias and ensure an equal number of data points for analysis across the entire range of GC content, promoters were evenly distributed into GC bins. As shown in Supplementary Fig. S2, there is a significant enrichment of active chromatin relative to nucleosomal fragments, regardless of the GC bin.

Furthermore, we observed a correlation between the cfDNAac fragments and markers of active chromatin structure. As an example, Fig. 2c shows a strong Pearson correlation (r = 0.75, p < 2.2e−308) between the cfDNAac fragment abundance and H3K4me1 narrow peak densities in 1 Mb (megabase) segments. However, the cfDNAnuc fragments do not show any notable correlation. Similar correlations are observed for other histone markers, e.g. H3K4me2 (r = 0.76, p < 2.2e−308) and H3K4me3 (r = 0.72, p < 2.2e−308; Supplementary Fig. S3a).

Comparison of the cfDNAac fragment distribution with other ENCODE categories also shows an enrichment of these fragments at insulator regions (Fig. 3a, Supplementary Data 1). Utilizing ENCODE ChromHMM annotations13, we found a strong enrichment of cfDNAac fragments across different chromatin states (Fig. 3b).Fig. 3 Regulatory-active chromatin fragments are enriched in regulatory regions.

a Genome browser view on a segment of chromosome 12 that represent a GM12878 insulator region predicted by GM12878. Top tracks are GM12878 ChIP-Seq –log(p-value) signals for either DNase-treated chromatin (first track) or selected histone modifications (H3K4me1, H3K4me2, H3K4me3, H3K27ac, and H3K36me3). The bottom two tracks represent enrichment of cfDNAnuc and cfDNAac genome-wide signals. Signals were calculated by log2(fold change) between the two fragment classes. b Comparison of the average signal distribution for different GM12878 chromatin states with a +/−10kb flank for cfDNAnuc (gray) and cfDNAac fragments (red). The y-axis scale for the cfDNAnuc and cfDNAac distributions represent modified coverage (see Methods).

Similar correlations are not observed with the cfDNAnuc fragments. Our results do indicate that cfDNAnuc fragments can footprint CCCTC-binding factor (CTCF19; Fig. 4, Supplementary Data 1), albeit with only a moderate correlation (r = 0.53, p = 2.3e−153). In contrast, the cfDNAac fragments showed a significantly higher correlation (r = 0.82, p < 2.2e−308) with CTCF peaks. Regulatory-active chromatin fragments also show a significant correlation with H3K27ac (Fig. 4; r = 0.73, p < 2.2e−308)— a histone modification known to be associated with enhancers— and with POLR2A (r = 0.67, p < 2.2e−308), which encodes for the largest subunit of RNA polymerase II. Similarly, other regions associated with enhancers20,21 (e.g. H2A.Z, r = 0.68, p < 2.2e−308) and EP300 narrow peak densities (r = 0.51, p = 3.2e−147; Supplementary Fig. S3b) also show positive correlations with cfDNAac fragment levels. We also evaluated whether cfDNAac and cfDNAnuc correlate with GM12878 DNase-I and ATAC-Seq signals (Supplementary Fig. S4). As expected, the number of cfDNAac fragments strongly correlate with both DNase-I (r = 0.75, p < 2.2e−308) and ATAC-Seq peak densities (r = 0.80, p < 2.2e−308), whereas cfDNAnuc did not correlate with either.Fig. 4 Regulatory-active chromatin fragments correlate with ChIP-Seq peak densities.

Pearson correlation between number of cfDNAnuc (gray) or cfDNAac fragments (red) to density of CTCF, H3K27ac, and POLR2A narrow peaks across the genome in 1 Mb segments. cfDNAac and cfDNAnuc fragments are based on the median of all sequenced plasma samples (n = 70).

Profiling of cfDNAac at promoter regions enables detection of circadian pattern

Circadian genes and their related transcription factors have been shown to regulate numerous physiological and immune processes22,23. Since there is a strong correlation of cfDNAac with histone modifications— which are indicative of active promoters— we wanted to determine if cfDNAac can be used to detect the diurnal patterns of expressed circadian genes found in whole blood. We focused our analysis on known circadian-specific markers with varying expression levels in whole blood (ranging from 1.61 to 1300 transcripts per million (TPM), with a median of 187.73), as defined by the Genotype-Tissue Expression Project24 (GTEx; Fig. 5, Supplementary Data 1).Fig. 5 Circadian patterns of highly expressed neutrophil genes.

a cfDNAac signal (y-axis) of CXCR4, CRY1 (top), CXCR2, PER1 (middle), CD62L, and TIMELESS (bottom) across time (x-axis). Signal is presented as the mean +/− SEM (standard error of the mean) and p-values correspond to the significance of rhythmicity evaluated by the zero-amplitude test using a period of 24 h (n = 5). b cfDNAac signal (y-axis) of CXCR4, CRY1 (top), CXCR2, PER1 (middle), CD62L, and TIMELESS (bottom) at different diurnal times (x-axis). p-values correspond to a paired two-sided t-test between the diurnal times (n = 5).

Among the three neutrophil-specific circadian markers— CXCR4 (p = 0.047), CXCR2 (p = 0.003), and CD62L (p < 0.001)— the abundance of the cfDNAac signal at promoters shows significant circadian oscillations (Fig. 5a). CXCR4 has been shown to have opposing oscillation to CXCR2 and CD62L22,25; and consistent with this observation, CXCR4 shows a decrease in cfDNAac levels while both CXCR2 and CD62L show higher signal abundance early in the morning. Additionally, the difference in cfDNAac signal was statistically significant between early morning and evening for CXCR4 (p = 0.038), CXCR2 (p = 0.026), and CD62L (p = 0.009; Fig. 5b). The cfDNAac signal for three known regulators of circadian rhythm— CRY1 (p = 0.009), TIMELESS (p = 0.007), and PER1 (p = 0.011)— also showed significant variations between early morning and evening. Consistent with their known interaction26, CRY1, PER1, and TIMELESS exhibit similar circadian patterns. However, statistical significance was observed only for CRY1 (p = 0.011) and TIMELESS (p = 0.019).

cfDNAac fragments correlate with RNA Pol II and are indicative of gene expression

Given that cfDNAac abundance can capture the transcriptional oscillations in expressed circadian-related genes, coupled with the observed correlations with regulatory elements mapped in ENCODE, we wanted to determine if the cfDNAac fragments could be used to measure gene expression at the transcriptional level. As previously determined in Fig. 4, the cfDNAac— and not the cfDNAnuc— fragment signal correlated with POLR2A ChIP-Seq densities. We therefore hypothesized that the cfDNAac fragments associated with RNA polymerase II should correlate with the transcriptional expression levels of specific genes.

Utilizing expression levels defined by GTEx24, we compared the read-depth patterns of the cfDNAnuc and cfDNAac fragments across gene bodies stratified by highest (top 10%) and lowest expressed (bottom 10%) genes. The cfDNAnuc signals show a coverage depletion around the TSS for highly expressed genes (Fig. 6a, Supplementary Data 1). As previously described27, this is likely due to less dense nucleosome packaging. However, cfDNAac fragments show increased signals for highly expressed genes, both around the TSS and across the entire gene body (Fig. 6a), as these fragments likely represent transcription factor-bound regions. This signal profile is consistent with promoter-proximal RNA polymerase II pausing28.Fig. 6 Regulatory-active chromatin fragments are informative of gene expression.

a Fragment distribution of cfDNAnuc and cfDNAac across the gene body. Genes from Gencodev42 were stratified into low expressed (gray) and high expressed (red) genes by filtering for the top and bottom 10% expressed genes based on GTEx whole blood median TPM. The genes were scaled to the same length from the transcription start site (TSS) to the transcription end site (TES) with 500 bp unscaled before and after the TES and TSS, respectively. b cfDNAac signal across the gene body from healthy plasma (y-axis) compared to gene expression levels in whole blood defined by GTEx (x-axis). Ten genes with similar expression values were aggregated into “metagenes”. c cfDNAac signal at promoters defined by Functional Annotation of the Mammalian Genome (FANTOM5) p1 sites (y-axis) compared to gene expression levels in whole blood defined by GTEx (x-axis). d cfDNAac signal across the gene bodies (y-axis) compared to cfDNAac signal at promoters (x-axis). cfDNAac and cfDNAnuc signals are based on the average of all sequenced plasma samples (n = 70).

Promoter and gene body scores should correlate with gene expression levels. To assess this, whole blood RNA-sequencing expression levels were obtained from GTEx24 and compared to promoter and gene body scores obtained from cfDNAac. These cfDNAac fragment scores across both gene bodies (Fig. 6b, r = 0.95, p < 2.2e−308) and promoters (Fig. 6c, r = 0.89, p < 2.2e−308) highly correlate with GTEx whole blood TPM values. The signal at promoters exhibits a plateau for highly expressed genes (Fig. 6c), likely due to a coverage depletion in the nucleosome-depleted region. However, the gene body cfDNAac signal does not exhibit this plateau. The promoter and gene body cfDNAac signals at both regions are significantly correlated with one another (Fig. 6d, r = 0.85, p < 2.2e−308), consistent with the expectation that the promoter and gene body are each bound by regulatory proteins and therefore are independently associated with gene expression. Additionally, we assessed the correlation by GC content, with each bin representing the 25th percentile of the average GC content of the gene body. The correlation with GTEx whole blood gene expression remains consistent across all GC bins (Supplementary Fig. S5).

Discussion

This study presents a novel approach for the enrichment and characterization of regulatory-active chromatin cfDNA fragments, or cfDNAac. Starting with an off-the-shelf kit, we modified the cfDNA extraction to enrich for cfDNAac by adjusting the amounts of lysis and binding buffers, which consequently changed the salt and detergent concentrations. We speculate that the changes likely induced conformational changes that facilitated and/or disrupted the interactions of DNA with regulatory proteins. Furthermore, the change in salt concentration likely affects the intermolecular interactions between DNA and the nanoparticles as previously described29,30. These changes in the extraction conditions significantly enriched the cfDNAac population. Compared to the off-the-shelf cfDNA capture kit, our active chromatin workflow yielded a 3.9- to 5.5-fold increase in the recovery of longer (>200 bp and >400 bp) cfDNA fragments (Supplementary Fig. S6). The enriched cfDNAac fragments are derived from regions of regulatory-active chromatin— gene promoters, enhancers, and transcription start sites— as well as gene bodies. Correlations with distinct diurnal expression patterns of circadian genes and with RNA polymerase II pausing profiles and gene expression levels, using both gene body and promoter scores, confirm the utility of cfDNAac fragments in reflecting active transcriptional regulation and gene expression.

Previous and current studies mapping the epigenome have largely relied on techniques such as ChIP-seq and/or ATAC-seq10,11. In fact, the open chromatin region, which is a hallmark of DNA regulatory elements, has been comprehensively profiled by ChIP-seq and ATAC-seq across different pathological and physiological conditions in cells. Specifically, ChIP-seq identifies global binding sites for regulatory proteins through DNA interaction. However, its dependence on protein-specific antibodies for immunoprecipitation prior to sequencing can be time-consuming and resource-intensive, limiting its use for genome-wide analysis of chromatin structure and transcriptional regulation10. Conversely, ATAC-seq can reveal genome-wide chromatin patterns, providing valuable insights into chromatin structure and gene expression. Unfortunately, ATAC-seq requires intact cells to capture the dynamic landscape of chromatin accessibility11, which is essential for understanding gene regulation, making this technique unfeasible for studies involving cfDNA.

A recent study by Baca et al. reported an alternative approach to extracting out active chromatin fragments by using ChIP-seq of cfDNA (cfChIP-seq)31. This method used antibodies targeting H3K4me3 and H3K27ac/panH3ac to selectively pull-down active promoters and enhancers, respectively. Consistent with our results (Fig. 6), the plasma signals derived from the immunoprecipitation of H3K4me3 at promoters in healthy subjects strongly correlated with GTEx gene expression (r = 0.95). However, while cfChIP-seq has provided valuable insights into specific gene regulation31, its dependence on antibodies diminishes the ability to identify unknown or novel chromatin modifications and potentially limit its detection capabilities. In contrast, the approach we have presented in this study is antibody independent. Our assay leverages a single analyte from a single assay to capture multiple epigenetic states— including multiple methylation and acetylation signals. Moreover, the fragment counts that map onto the gene bodies can be used to reconstruct gene expression profiles (see Fig. 6b), thus offering more comprehensive biological insights that extend beyond the traditional regulatory regions such as promoters and enhancers.

Recent advancements in cfDNA fragmentomics have enhanced the ability to infer gene regulation and tissue of origin of cfDNA utilizing end-motifs and motif diversity scores32–35; however, their ability to capture the molecular complexity and heterogeneity of disease biology remains limited. Similarly, analyses of large-scale cfDNA fragmentation patterns36 are challenged by the limited ability to dissect and link specific gene-regulatory aberrations in disease conditions, since most dysregulated genes are likely found in more active chromatin regions rather than highly organized nucleosome regions. One attempt by Zhou et al. focused on cfDNA “fragmentation hotspots” instead of identifying nucleosome-occupied regions, or “fragmentation cold-spots”37. The researchers used a computational approach to enrich for gene-regulatory elements and open chromatin regions. In contrast to the in-silico approach, our technology optimizes the extraction of active chromatin fragments to increase the signal-to-noise ratio. Both methods offer novel insights into gene-regulatory mechanisms, which could enhance the detection of nuanced pathological conditions. Combining these approaches could further enhance signal-to-noise and thus warrants further exploration in future iterations.

The analysis of cfDNAac is disease-agnostic and, as such, can be applied across a range of indications, including chronic inflammation, autoimmune conditions, and cancer. Given the regulatory role of epigenetic modifications on gene expression and its ability to influence the development and progression of pathological conditions38–40, detecting these epigenetic changes through the use of cfDNAac presents a new opportunity for early disease detection and intervention. Additionally, leveraging cfDNAac for gene expression analysis could provide valuable molecular insights into dysregulated signaling pathways and disease drivers. In fact, preliminary data have shown the ability of cfDNAac analysis in identifying synovium-specific gene expression signatures in the blood of patients with rheumatoid arthritis by detecting both organ-specific and systemic biological signals41.

In summary, our study highlights the unique use of active chromatin cfDNA analysis to further advance our understanding of gene regulation and expression in different physiological conditions. The ability of this approach to deliver data comparable to ChIP- and ATAC-seq-derived information and reconstruct gene expression profiles opens up new avenues in precision medicine and molecular diagnostics. Given that the half-life of cfDNA in blood is a couple of hours8, analyzing cfDNAac can provide real-time data on the status of contributing tissues. This non-invasive method offers an opportunity to monitor dynamic changes in epigenetic and gene expression patterns, paving the way for a wide array of clinical applications in disease diagnosis and management.

Methods

Sample collection

Altasciences, a contract research organization based in Montreal, Canada, recruited twenty self-reported healthy individuals for this study. The recruitment process was carried out in accordance with the guidelines of Good Clinical Practice (GCP). The study was approved by an Advarra IRB/Ethics Committee and Health Canada (Aurora, ON, Canada), and only participants who provided written informed consent were included in the present study. All ethical regulations relevant to human research participants were followed.

For each study participant, a blood specimen was collected every 4 hours for about 2.5 days and collected in two STRECK Cell-Free DNA blood collection tubes (20 mL total) following the manufacturer’s protocol (Supplementary Table S1). Ultimately, blood samples from fourteen time points were obtained from each participant. The tubes were shipped overnight to Aqtual, Inc. (Hayward, CA). The blood specimen tubes were put in a refrigerated centrifuge and spun for 20 min at 300 × g. After the first spin, the plasma layer was transferred to a sterile 15-mL conical tube and centrifuged again at 5000 × g for 10 min. The resulting plasma was aliquoted into 2-mL sterile, cryogenic screw-cap tubes and stored at –80 °C.

cfDNAac extraction and sequencing library preparation

The Apostle kit (Apostle Bio, Pleasanton, CA) has been used in many studies as a standard cfDNA extraction workflow42–46. The Apostle MiniMax High Efficiency cfDNA Isolation Kit was used as the off-the-shelf extraction kit because of its superior performance across a range of input volumes and ability to recover high quantities of cfDNA from different biofluids47–49. Modifications to the Apostle kit were made for our cfDNAac-enriched capture workflow. To optimize for cfDNAac capture, we tested different volumes of lysis (60–200 μL) and binding buffers (800–2000 μL) while maintaining a fixed plasma input volume of 2 mL. Optimal capture of active chromatin cfDNA fragments was achieved with a final volume of 100 μL and 1250 μL for lysis and binding buffers, respectively. The sample lysis step was modified by heating it at 58 °C for 30 min, rather than 60 °C for 20 min.

In both the off-the-shelf and modified workflows, magnetic nanoparticles provided by the commercial kit (Apostle) were used; per the certificate of analysis for the batch of nanoparticles used in this study, the diameter of the particles ranged from 200–600 nm. The nanoparticles were added to the plasma samples and incubated in the presence of binding buffer for about 30 min. To prevent the loss of fragments during processing, samples were not transferred to a new tube following the first wash. Fragments were eluted using 32 μL of elution buffer, and 30 μL of the eluted fragments were transferred to a new 1.5-mL tube for subsequent library preparation. The cfDNA was quantified using the Qubit dsDNA High Sensitivity Assay (Thermo Scientific, Waltham, MA), and fragment lengths were determined using the TapeStation 4200 High Sensitivity D5000 Assay (Agilent, Santa Clara, CA) following manufacturers’ instructions (Supplementary Fig. S7a).

Libraries for sequencing were prepared from the total cfDNA extracted using the Kapa Biosystems, Inc. (Roche, Basel, Switzerland) Hyper Prep Kit according to manufacturer’s protocol. The post-PCR solid-phase reversible immobilization (SPRI) ratio for samples prepared with the active chromatin capture workflow was reduced to 0.7X, allowing for further enrichment of cfDNAac fragments. The libraries were quantified and diluted using the Qubit dsDNA High Sensitivity Kit (Thermo Scientific) and TapeStation High Sensitivity D5000 assay (Agilent) (Supplementary Fig. S7b). Fourteen sample libraries from a single individual were pooled together equally by mass and sequenced using an S4 200 cycle kit on a NovaSeq 6000 sequencer (Illumina, San Diego, CA) (Supplementary Fig. S7c).

Sequencing reads were aligned to the human genome (hg19) using Bowtie250 (v2.4.4) with ‘no-mixed’ and ‘no-discordant’ flags. The maximum fragment length for valid paired end alignments was modified to allow the inclusion of longer cell-free DNA fragments (-X 1500). Duplicate fragments were marked and removed using SAMtools markdup51 (v1.16), and reads that did not align uniquely (-q 40) were discarded. BAM files were converted to data frames containing fragment information and processed using R (v4.1.3) and Python (v3.9.13) scripts (see ‘Code availability’).

Deconvolution of cfDNA fragments: nucleosomal from regulatory-active chromatin cfDNA

To differentiate between cfDNAnuc and cfDNAac fragments, the fragment size distribution (ranging from 100 to 1500 base pairs) from a single timepoint of subject AS-001 was divided into 12 bins. Genome-wide signals (in bigWig format) of 5 individuals at all 14 timepoints (n = 70) were obtained using abundances of fragment midpoints across the entire genome in 30 bp windows. The mean fragment count profile was calculated from a single individual (AS-001, timepoint 1) across GM12878 ChromHMM active promoter regions for each of the 12 bins (Supplementary Fig. S7d). Unsupervised clustering of the average profile identified two main clusters, with 6 bins assigned in each cluster. The cfDNAnuc and cfDNAac signals were averaged across five individuals to create an overall profile. Log2-fold changes between cfDNAnuc and cfDNAac signals were calculated in 300 bp windows across the genome. These binned fold-changes were visually compared to processed ChIP-Seq data using IGV (v2.14.1; Figs. 2a and 3a). ChIP-seq and DNase I data were downloaded from the Roadmap Epigenomic consortium database 13 (E116), which is available through ENCODE. ATAC-Seq data was downloaded from ENCODE and lifted over to hg19.

cfDNAac fragment enrichment across regions in the genome

The deepTools252 computeMatrix was used to examine the signal distribution relative to TSS from EPDnew17 with a +/−10 kb flanking region (-a 10000 -b 10000) and a window size of 50 (-bs 50; Fig. 2b). In addition to TSS obtained from the EPDnew, signal profiles were computed for various ChromHMM states for the GM12878 cell line. The ChromHMM states were acquired from the UCSC Genome Browser and analyzed using the same parameters as those of the TSS locations (Fig. 3b).

To compare the correlations of cfDNAnuc and cfDNAac fragments to public ChIP-Seq peak densities, GM12878 narrow peak calls were downloaded from the Roadmap Epigenomic consortium database for H3K4me1, H3K4me2, H3K4me3, H3K27ac, H2A.Z, and DNase I. Narrow peak calls for other proteins, such as CTCF, POLR2A, and EP300, were downloaded directly from ENCODE. ATAC-Seq narrow peaks were downloaded from ENCODE and lifted over from hg38 to hg19 using UCSC’s liftover tool. For each of the five individuals, the total number of cfDNAnuc and cfDNAac fragments were calculated in 1 megabase (Mb) windows, and the medians of these fragment counts were correlated with the number of narrow peaks for each of the histone marks and transcription factors. 1 Mb windows were filtered to include only windows with at least one ChIP-Seq peak and 50,000 fragments. To determine the presence of cfDNAac fragments, each fragment was condensed to its midpoint before checking for overlaps with genomic features such as introns, CpG islands, CpG shores, CpG shelves, promoters, and repressed regions. Promoters were defined by EPDnew and repressed regions were identified using ChromHMM predictions from GM12878. The number of overlaps between cfDNAnuc and cfDNAac fragments and each genomic feature was divided by the total number of fragments (in millions; Supplementary Fig. S1).

Patterns of circadian-specific markers

To analyze the patterns of circadian-specific markers (e.g CXCR4, CXCR2, CD62L, CRY1, PER1, and TIMELESS), the cfDNAac abundance from the generated bigWigs was first calculated for each promoter, which was defined by the Functional Annotation of the Mammalian Genome (FANTOM5) project53 that overlapped the gene region. The promoter for each gene was chosen based on the minimum distance of the promoter to the most expressed exon in GTEx. For genes with similar expression throughout the exons, the p1 promoter was chosen. For each individual, timepoints across multiple days were collapsed into unique timepoints (12PM, 4PM, 8PM, 12AM, 4AM, and 8AM) using the median cfDNAac bigWig signal. The cfDNAac signal at each timepoint was defined as the median signal across 150 bp downstream of the selected promoter region for each gene. The mean and SEM were calculated using the five sequenced individuals (i.e. n = 5). Significance of rhythmicity was evaluated by the zero-amplitude test using a period of 24 using CosinorPy54. A paired t-test was used to evaluate significant differences between diurnal times.

Correlation of cfDNAac fragments with gene expression

To analyze the distribution of cfDNAnuc and cfDNAac signals throughout the gene body, we generated scaled gene profiles using deepTools2 scale-regions (-m 4000 -b 2000 -a 2000 –unscaled5prime 500 –unscaled3prime 500) with transcripts downloaded from Gencodev4255 (v42lift37). The signals were separated based on the median whole blood TPM (transcripts per million) values from GTEx and divided into the highest 10% and lowest 10% expressed genes.

The correlations between cfDNAac signals at both the promoters and gene bodies were compared with GTEx whole blood TPM annotations. To measure cfDNAac signals at promoters, the number of cfDNAac fragments and cfDNAnuc fragments within +/−500 base pairs of p1 promoters defined by FANTOM5 were counted and corrected for GC by fitting a LOWESS (locally weighted scatterplot smoothing) curve using a smoothing parameter of 0.l. The cfDNAac fragments were then normalized to the number of cfDNAnuc fragments to obtain a promoter score. For gene bodies, we created a gene model reference set using a collapse annotation python script from gtex-pipeline (https://github.com/broadinstitute/gtex-pipeline). The number of cfDNAac fragments and cfDNAnuc fragments at each exon were first scaled by the exon length and then corrected using the same procedure as with the promoter regions. Finally, the normalized counts were summed across all exons and scaled by the number of exons to obtain a cfDNAac gene score. When aggregating across multiple genes, genes were ranked according to their expression values from whole blood (in GTEx) and combined using approximately 10 genes as previously described56 (Fig. 6b).

Statistics and reproducibility

A standard threshold of p < 0.05 was used to determine significance for all hypothesis testing. A t-test was used to compare the enrichment of cfDNAac fragments with cfDNAnuc fragments. To analyze circadian patterns, a zero-amplitude test was used to determine significance of rhythmicity, and a paired t-test was used to evaluate the differences between diurnal times. Error bars were calculated with n ≥ 3 samples. Exact p-values were reported when possible; otherwise, p-values below the machine’s minimum floating value were reported as p < 2.2e−308. Huber regression was used to fit all regression lines. cfDNAac and cfDNAnuc bigWigs for a representative subject are available on Zenodo (see Data Availability) at 10.5281/zenodo.11266304. The source data for the main figures can be found in Supplementary Data 1.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Supplementary Information

Description of Additional Supplementary File

Supplementary Data 1

Reporting Summary

Supplementary information

The online version contains supplementary material available at 10.1038/s42003-024-06769-3.

Author contributions

D.A. developed the concept. D.A. and R.R. designed the study. K.D. and R.C. developed the chromatin capture method with the help of D.A. K.D., R.C., R.B., and K.T. performed the chromatin capture experiments. K.L. and K.T.L. analyzed the data with guidance from D.A. and R.R. K.L., R.R., M.C.L., and D.A. wrote the paper with inputs from all authors.

Peer review

Peer review information

Communications Biology thanks the anonymous reviewers for their contribution to the peer review of this work. Primary Handling Editor: Christina Karlsson Rosenthal.

Data availability

The whole genome sequencing (WGS) data is protected under GDPR and cannot be shared publicly. Accordingly, we have provided processed bigwig files which is available in Zenodo57: 10.5281/zenodo.11266304. The source data for the main figures in the paper can be found in Supplementary Data 1. The following public datasets were also used: Transcription start sites (https://epd.expasy.org/epd/human/human_database.php?db=human), processed GM12878 histone mark and DNaseI signal tracks (https://egg2.wustl.edu/roadmap/data/byFileType/signal/consolidated/macs2signal/pval/), processed GM12878 histone mark and DNaseI narrow peaks (https://egg2.wustl.edu/roadmap/data/byFileType/peaks/consolidated/narrowPeak/), processed GM12878 CTCF peaks (https://www.encodeproject.org/files/ENCFF710VEH/), processed GM12878 POLR2A peaks (https://www.encodeproject.org/files/ENCFF120VUT/), processed GM12878 EP300 peaks (https://www.encodeproject.org/files/ENCFF222XPQ/), processed GM12878 ATAC-Seq narrow peak and bigwig (https://www.encodeproject.org/experiments/ENCSR095QNB/), GM12878 ChromHMM (https://hgdownload.cse.ucsc.edu/goldenpath/hg19/encodeDCC/wgEncodeBroadHmm/wgEncodeBroadHmmGm12878HMM.bed.gz), GTExv8 whole blood TPM values (https://storage.cloud.google.com/adult-gtex/bulk-gex/v8/rna-seq/counts-by-tissue/gene_reads_2017-06-05_v8_whole_blood.gct.gz), Gencodev42 GTF, FANTOM5 promoters (https://fantom.gsc.riken.jp/5/datafiles/latest/extra/CAGE_peaks/hg19.cage_peak_phase1and2combined_ann.txt.gz), and CpGIslands (http://hgdownload.soe.ucsc.edu/goldenPath/hg19/database/cpgIslandExt.txt.gz).

Code availability

Code for processing sequencing data is available on Zenodo57 at 10.5281/zenodo.11266304.

Competing interests

D.A. and R.R. are co-founders and shareholders of Aqtual, Inc. All other authors are shareholders of Aqtual, Inc.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Mandel P Metais P Nuclear Acids In Human Blood Plasma C. R. Seances Soc. Biol. Fil. 1948 142 241 243 18875018
Mandel, P. & Metais, P. Nuclear Acids In Human Blood Plasma. C. R. Seances Soc. Biol. Fil. 142, 241–243 (1948).18875018
2. Tan EM Schur PH Carr RI Kunkel HG Deoxybonucleic acid (DNA) and antibodies to DNA in the serum of patients with systemic lupus erythematosus J. Clin. Invest. 1966 45 1732 1740 10.1172/JCI105479 4959277
Tan, E. M., Schur, P. H., Carr, R. I. & Kunkel, H. G. Deoxybonucleic acid (DNA) and antibodies to DNA in the serum of patients with systemic lupus erythematosus. J. Clin. Invest. 45, 1732–1740 (1966).4959277 10.1172/JCI105479
3. Laudat MH The hormonal state of pregnancy: modification of cortisol and testosterone Ann. Endocrinol. 1987 48 334 338
Laudat, M. H. et al. The hormonal state of pregnancy: modification of cortisol and testosterone. Ann. Endocrinol. 48, 334–338 (1987).
4. Stroun M Anker P Lyautey J Lederrey C Maurice PA Isolation and characterization of DNA from the plasma of cancer patients Eur. J. Cancer Clin. Oncol. 1987 23 707 712 10.1016/0277-5379(87)90266-5 3653190
Stroun, M., Anker, P., Lyautey, J., Lederrey, C. & Maurice, P. A. Isolation and characterization of DNA from the plasma of cancer patients. Eur. J. Cancer Clin. Oncol. 23, 707–712 (1987).3653190 10.1016/0277-5379(87)90266-5
5. Heitzer E Auinger L Speicher MR Cell-Free DNA and Apoptosis: How Dead Cells Inform About the Living Trends Mol. Med. 2020 26 519 528 10.1016/j.molmed.2020.01.012 32359482
Heitzer, E., Auinger, L. & Speicher, M. R. Cell-Free DNA and Apoptosis: How Dead Cells Inform About the Living. Trends Mol. Med. 26, 519–528 (2020).32359482 10.1016/j.molmed.2020.01.012
6. Hu Z Chen H Long Y Li P Gu Y The main sources of circulating cell-free DNA: Apoptosis, necrosis and active secretion Crit. Rev. Oncol. Hematol. 2021 157 103166 10.1016/j.critrevonc.2020.103166 33254039
Hu, Z., Chen, H., Long, Y., Li, P. & Gu, Y. The main sources of circulating cell-free DNA: Apoptosis, necrosis and active secretion. Crit. Rev. Oncol. Hematol. 157, 103166 (2021).33254039 10.1016/j.critrevonc.2020.103166
7. Cheng SH Noninvasive prenatal testing by nanopore sequencing of maternal plasma DNA: feasibility assessment Clin. Chem. 2015 61 1305 1306 10.1373/clinchem.2015.245076 26286915
Cheng, S. H. et al. Noninvasive prenatal testing by nanopore sequencing of maternal plasma DNA: feasibility assessment. Clin. Chem. 61, 1305–1306 (2015).26286915 10.1373/clinchem.2015.245076
8. Kustanovich A Schwartz R Peretz T Grinshpun A Life and death of circulating cell-free DNA Cancer Biol. Ther. 2019 20 1057 1067 10.1080/15384047.2019.1598759 30990132
Kustanovich, A., Schwartz, R., Peretz, T. & Grinshpun, A. Life and death of circulating cell-free DNA. Cancer Biol. Ther. 20, 1057–1067 (2019).30990132 10.1080/15384047.2019.1598759
9. Noll M Subunit structure of chromatin Nature 1974 251 249 251 10.1038/251249a0 4422492
Noll, M. Subunit structure of chromatin. Nature 251, 249–251 (1974).4422492 10.1038/251249a0
10. Park PJ ChIP-seq: advantages and challenges of a maturing technology Nat. Rev. Genet 2009 10 669 680 10.1038/nrg2641 19736561
Park, P. J. ChIP-seq: advantages and challenges of a maturing technology. Nat. Rev. Genet 10, 669–680 (2009).19736561 10.1038/nrg2641
11. Grandi FC Modi H Kampman L Corces MR Chromatin accessibility profiling by ATAC-seq Nat. Protoc. 2022 17 1518 1552 10.1038/s41596-022-00692-9 35478247
Grandi, F. C., Modi, H., Kampman, L. & Corces, M. R. Chromatin accessibility profiling by ATAC-seq. Nat. Protoc. 17, 1518–1552 (2022).35478247 10.1038/s41596-022-00692-9
12. Leest, P. V. et al. Comparison of Circulating Cell-Free DNA Extraction Methods for Downstream Analysis in Cancer Patients. Cancers 12, 10.3390/cancers12051222 (2020).
13. Consortium EP An integrated encyclopedia of DNA elements in the human genome Nature 2012 489 57 74 10.1038/nature11247 22955616
Consortium, E. P. An integrated encyclopedia of DNA elements in the human genome. Nature 489, 57–74 (2012).22955616 10.1038/nature11247
14. Kornberg RD Klug A The nucleosome Sci. Am. 1981 244 52 64 10.1038/scientificamerican0281-52 7209486
Kornberg, R. D. & Klug, A. The nucleosome. Sci. Am. 244, 52–64 (1981).7209486 10.1038/scientificamerican0281-52
15. Roadmap Epigenomics C Integrative analysis of 111 reference human epigenomes Nature 2015 518 317 330 10.1038/nature14248 25693563
Roadmap Epigenomics, C. et al. Integrative analysis of 111 reference human epigenomes. Nature 518, 317–330 (2015).25693563 10.1038/nature14248
16. Consortium EP Expanded encyclopaedias of DNA elements in the human and mouse genomes Nature 2020 583 699 710 10.1038/s41586-020-2493-4 32728249
Consortium, E. P. et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature 583, 699–710 (2020).32728249 10.1038/s41586-020-2493-4
17. Meylan P Dreos R Ambrosini G Groux R Bucher P EPD in 2020: enhanced data visualization and extension to ncRNA promoters Nucleic Acids Res. 2020 48 D65 D69 31680159
Meylan, P., Dreos, R., Ambrosini, G., Groux, R. & Bucher, P. EPD in 2020: enhanced data visualization and extension to ncRNA promoters. Nucleic Acids Res. 48, D65–D69 (2020).31680159
18. Gates LA Foulds CE O’Malley BW Histone Marks in the ‘Driver’s Seat’: Functional Roles in Steering the Transcription Cycle Trends Biochem. Sci. 2017 42 977 989 10.1016/j.tibs.2017.10.004 29122461
Gates, L. A., Foulds, C. E. & O’Malley, B. W. Histone Marks in the ‘Driver’s Seat’: Functional Roles in Steering the Transcription Cycle. Trends Biochem. Sci. 42, 977–989 (2017).29122461 10.1016/j.tibs.2017.10.004
19. Snyder MW Kircher M Hill AJ Daza RM Shendure J Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin Cell 2016 164 57 68 10.1016/j.cell.2015.11.050 26771485
Snyder, M. W., Kircher, M., Hill, A. J., Daza, R. M. & Shendure, J. Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. Cell 164, 57–68 (2016).26771485 10.1016/j.cell.2015.11.050
20. Brunelle M The histone variant H2A.Z is an important regulator of enhancer activity Nucleic Acids Res. 2015 43 9742 9756 26319018
Brunelle, M. et al. The histone variant H2A.Z is an important regulator of enhancer activity. Nucleic Acids Res. 43, 9742–9756 (2015).26319018
21. Zhou, P. et al. Mapping cell type-specific transcriptional enhancers using high affinity, lineage-specific Ep300 bioChIP-seq. Elife 6, 10.7554/eLife.22039 (2017).
22. Ovadia S Ozcan A Hidalgo A The circadian neutrophil, inside-out J. Leukoc. Biol. 2023 113 555 566 10.1093/jleuko/qiad038 36999376
Ovadia, S., Ozcan, A. & Hidalgo, A. The circadian neutrophil, inside-out. J. Leukoc. Biol. 113, 555–566 (2023).36999376 10.1093/jleuko/qiad038
23. Palomino-Segura, M. & Hidalgo, A. Circadian immune circuits. J. Exp. Med. 218, 10.1084/jem.20200798 (2021).
24. Consortium GT Genetic effects on gene expression across human tissues Nature 2017 550 204 213 10.1038/nature24277 29022597
Consortium, G. T. et al. Genetic effects on gene expression across human tissues. Nature 550, 204–213 (2017).29022597 10.1038/nature24277
25. Adrover JM A Neutrophil Timer Coordinates Immune Defense and Vascular Protection Immunity 2019 50 390 402 e310 10.1016/j.immuni.2019.01.002 30709741
Adrover, J. M. et al. A Neutrophil Timer Coordinates Immune Defense and Vascular Protection. Immunity 50, 390–402 e310 (2019).30709741 10.1016/j.immuni.2019.01.002
26. Bunney WE Bunney BG Molecular clock genes in man and lower animals: possible implications for circadian abnormalities in depression Neuropsychopharmacology 2000 22 335 345 10.1016/S0893-133X(99)00145-1 10700653
Bunney, W. E. & Bunney, B. G. Molecular clock genes in man and lower animals: possible implications for circadian abnormalities in depression. Neuropsychopharmacology 22, 335–345 (2000).10700653 10.1016/S0893-133X(99)00145-1
27. Ulz P Inferring expressed genes by whole-genome sequencing of plasma DNA Nat. Genet 2016 48 1273 1278 10.1038/ng.3648 27571261
Ulz, P. et al. Inferring expressed genes by whole-genome sequencing of plasma DNA. Nat. Genet 48, 1273–1278 (2016).27571261 10.1038/ng.3648
28. Yu M RNA polymerase II-associated factor 1 regulates the release and phosphorylation of paused RNA polymerase II Science 2015 350 1383 1386 10.1126/science.aad2338 26659056
Yu, M. et al. RNA polymerase II-associated factor 1 regulates the release and phosphorylation of paused RNA polymerase II. Science 350, 1383–1386 (2015).26659056 10.1126/science.aad2338
29. Shi B Shin YK Hassanali AA Singer SJ DNA Binding to the Silica Surface J. Phys. Chem. B 2015 119 11030 11040 10.1021/acs.jpcb.5b01983 25966319
Shi, B., Shin, Y. K., Hassanali, A. A. & Singer, S. J. DNA Binding to the Silica Surface. J. Phys. Chem. B 119, 11030–11040 (2015).25966319 10.1021/acs.jpcb.5b01983
30. Vandeventer PE Multiphasic DNA adsorption to silica surfaces under varying buffer, pH, and ionic strength conditions J. Phys. Chem. B 2012 116 5661 5670 10.1021/jp3017776 22537288
Vandeventer, P. E. et al. Multiphasic DNA adsorption to silica surfaces under varying buffer, pH, and ionic strength conditions. J. Phys. Chem. B 116, 5661–5670 (2012).22537288 10.1021/jp3017776
31. Baca SC Liquid biopsy epigenomic profiling for cancer subtyping Nat. Med. 2023 29 2737 2741 10.1038/s41591-023-02605-z 37865722
Baca, S. C. et al. Liquid biopsy epigenomic profiling for cancer subtyping. Nat. Med. 29, 2737–2741 (2023).37865722 10.1038/s41591-023-02605-z
32. Jiang P Plasma DNA End-Motif Profiling as a Fragmentomic Marker in Cancer, Pregnancy, and Transplantation Cancer Discov. 2020 10 664 673 10.1158/2159-8290.CD-19-0622 32111602
Jiang, P. et al. Plasma DNA End-Motif Profiling as a Fragmentomic Marker in Cancer, Pregnancy, and Transplantation. Cancer Discov. 10, 664–673 (2020).32111602 10.1158/2159-8290.CD-19-0622
33. An Y DNA methylation analysis explores the molecular basis of plasma cell-free DNA fragmentation Nat. Commun. 2023 14 287 10.1038/s41467-023-35959-6 36653380
An, Y. et al. DNA methylation analysis explores the molecular basis of plasma cell-free DNA fragmentation. Nat. Commun. 14, 287 (2023).36653380 10.1038/s41467-023-35959-6
34. Sun K Orientation-aware plasma cell-free DNA fragmentation analysis in open chromatin regions informs tissue of origin Genome Res. 2019 29 418 427 10.1101/gr.242719.118 30808726
Sun, K. et al. Orientation-aware plasma cell-free DNA fragmentation analysis in open chromatin regions informs tissue of origin. Genome Res. 29, 418–427 (2019).30808726 10.1101/gr.242719.118
35. Lo, Y. M. D., Han, D. S. C., Jiang, P. & Chiu, R. W. K. Epigenetics, fragmentomics, and topology of cell-free DNA in liquid biopsies. Science 372, 10.1126/science.aaw3616 (2021).
36. Cristiano S Genome-wide cell-free DNA fragmentation in patients with cancer Nature 2019 570 385 389 10.1038/s41586-019-1272-6 31142840
Cristiano, S. et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 570, 385–389 (2019).31142840 10.1038/s41586-019-1272-6
37. Zhou X CRAG: de novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing Genome Med. 2022 14 138 10.1186/s13073-022-01141-8 36482487
Zhou, X. et al. CRAG: de novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing. Genome Med. 14, 138 (2022).36482487 10.1186/s13073-022-01141-8
38. Wu H Chen Y Zhu H Zhao M Lu Q The Pathogenic Role of Dysregulated Epigenetic Modifications in Autoimmune Diseases Front. Immunol. 2019 10 2305 10.3389/fimmu.2019.02305 31611879
Wu, H., Chen, Y., Zhu, H., Zhao, M. & Lu, Q. The Pathogenic Role of Dysregulated Epigenetic Modifications in Autoimmune Diseases. Front. Immunol. 10, 2305 (2019).31611879 10.3389/fimmu.2019.02305
39. Xu X Metabolic reprogramming and epigenetic modifications in cancer: from the impacts and mechanisms to the treatment potential Exp. Mol. Med. 2023 55 1357 1370 10.1038/s12276-023-01020-1 37394582
Xu, X. et al. Metabolic reprogramming and epigenetic modifications in cancer: from the impacts and mechanisms to the treatment potential. Exp. Mol. Med. 55, 1357–1370 (2023).37394582 10.1038/s12276-023-01020-1
40. Yang C Epigenetic Regulation in the Pathogenesis of Rheumatoid Arthritis Front Immunol. 2022 13 859400 10.3389/fimmu.2022.859400 35401513
Yang, C. et al. Epigenetic Regulation in the Pathogenesis of Rheumatoid Arthritis. Front Immunol. 13, 859400 (2022).35401513 10.3389/fimmu.2022.859400
41. Taylor, P. C. et al. Detection of Synovial Signatures in Peripheral Blood of Patients With Rheumatoid Arthritis via a Novel Blood-Based DNA Capture Assay. in American College of Rheumatology (ACR) Convergence Conference (ACR, 2023).
42. Bae M Integrative modeling of tumor genomes and epigenomes for enhanced cancer diagnosis by cell-free DNA Nat. Commun. 2023 14 2017 10.1038/s41467-023-37768-3 37037826
Bae, M. et al. Integrative modeling of tumor genomes and epigenomes for enhanced cancer diagnosis by cell-free DNA. Nat. Commun. 14, 2017 (2023).37037826 10.1038/s41467-023-37768-3
43. Heeke S Tumor- and circulating-free DNA methylation identifies clinically relevant small cell lung cancer subtypes Cancer Cell 2024 42 225 237 e225 10.1016/j.ccell.2024.01.001 38278149
Heeke, S. et al. Tumor- and circulating-free DNA methylation identifies clinically relevant small cell lung cancer subtypes. Cancer Cell 42, 225–237 e225 (2024).38278149 10.1016/j.ccell.2024.01.001
44. Jin, S. et al. Efficient detection and post-surgical monitoring of colon cancer with a multi-marker DNA methylation liquid biopsy. Proc Natl Acad Sci USA 118, 10.1073/pnas.2017421118 (2021).
45. Palmer CD Individualized, heterologous chimpanzee adenovirus and self-amplifying mRNA neoantigen vaccine for advanced metastatic solid tumors: phase 1 trial interim results Nat. Med. 2022 28 1619 1629 10.1038/s41591-022-01937-6 35970920
Palmer, C. D. et al. Individualized, heterologous chimpanzee adenovirus and self-amplifying mRNA neoantigen vaccine for advanced metastatic solid tumors: phase 1 trial interim results. Nat. Med. 28, 1619–1629 (2022).35970920 10.1038/s41591-022-01937-6
46. Wang P Simultaneous analysis of mutations and methylations in circulating cell-free DNA for hepatocellular carcinoma detection Sci. Transl. Med. 2022 14 eabp8704 10.1126/scitranslmed.abp8704 36417488
Wang, P. et al. Simultaneous analysis of mutations and methylations in circulating cell-free DNA for hepatocellular carcinoma detection. Sci. Transl. Med. 14, eabp8704 (2022).36417488 10.1126/scitranslmed.abp8704
47. Stark A Pisanic TR 2nd Herman JG Wang TH High-throughput sample processing for methylation analysis in an automated, enclosed environment SLAS Technol. 2022 27 172 179 10.1016/j.slast.2021.12.002 35058199
Stark, A., Pisanic, T. R. 2nd, Herman, J. G. & Wang, T. H. High-throughput sample processing for methylation analysis in an automated, enclosed environment. SLAS Technol. 27, 172–179 (2022).35058199 10.1016/j.slast.2021.12.002
48. Prasad AR Optimization of high-volume cell free DNA extraction and end-to-end automation for TSO 500 ctDNA library prep Cancer Res. 2024 84 2298 10.1158/1538-7445.AM2024-2298
Prasad, A. R. et al. Optimization of high-volume cell free DNA extraction and end-to-end automation for TSO 500 ctDNA library prep. Cancer Res. 84, 2298 (2024).10.1158/1538-7445.AM2024-2298
49. Roseman RP Improved conversion in extraction, library construction, and capture improve sensitivity for variants in liquid biopsy samples Cancer Res. 2020 80 5863 10.1158/1538-7445.AM2020-5863
Roseman, R. P. et al. Improved conversion in extraction, library construction, and capture improve sensitivity for variants in liquid biopsy samples. Cancer Res. 80, 5863 (2020).10.1158/1538-7445.AM2020-5863
50. Langmead B Salzberg SL Fast gapped-read alignment with Bowtie 2 Nat. Methods 2012 9 357 359 10.1038/nmeth.1923 22388286
Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012).22388286 10.1038/nmeth.1923
51. Li H The Sequence Alignment/Map format and SAMtools Bioinformatics 2009 25 2078 2079 10.1093/bioinformatics/btp352 19505943
Li, H. et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25, 2078–2079 (2009).19505943 10.1093/bioinformatics/btp352
52. Ramirez F deepTools2: a next generation web server for deep-sequencing data analysis Nucleic Acids Res. 2016 44 W160 W165 10.1093/nar/gkw257 27079975
Ramirez, F. et al. deepTools2: a next generation web server for deep-sequencing data analysis. Nucleic Acids Res. 44, W160–W165 (2016).27079975 10.1093/nar/gkw257
53. Consortium F A promoter-level mammalian expression atlas Nature 2014 507 462 470 10.1038/nature13182 24670764
Consortium, F. et al. A promoter-level mammalian expression atlas. Nature 507, 462–470 (2014).24670764 10.1038/nature13182
54. Moskon M CosinorPy: a python package for cosinor-based rhythmometry BMC Bioinforma. 2020 21 485 10.1186/s12859-020-03830-w
Moskon, M. CosinorPy: a python package for cosinor-based rhythmometry. BMC Bioinforma. 21, 485 (2020).10.1186/s12859-020-03830-w
55. Frankish A Gencode 2021 Nucleic Acids Res. 2021 49 D916 D923 10.1093/nar/gkaa1087 33270111
Frankish, A. et al. Gencode 2021. Nucleic Acids Res. 49, D916–D923 (2021).33270111 10.1093/nar/gkaa1087
56. Esfahani MS Inferring gene expression from cell-free DNA fragmentation profiles Nat. Biotechnol. 2022 40 585 597 10.1038/s41587-022-01222-4 35361996
Esfahani, M. S. et al. Inferring gene expression from cell-free DNA fragmentation profiles. Nat. Biotechnol. 40, 585–597 (2022).35361996 10.1038/s41587-022-01222-4
57. Lai, K. et al. Extracting Regulatory Chromatin Footprint form Cell-Free DNA. Zenodo10.5281/zenodo.11266304 (2024).
