
==== Front
J Phys Chem B
J Phys Chem B
jp
jpcbfk
The Journal of Physical Chemistry. B
1520-6106
1520-5207
American Chemical Society

37681731
10.1021/acs.jpcb.3c03538
Article
Structure and Dynamics of DNA and RNA Double Helices Formed by d(CTG), d(GTC), r(CUG), and r(GUC) Trinucleotide Repeats and Associated DNA–RNA Hybrids
Fakharzadeh Ashkan †§
Qu Jing †§
Pan Feng ‡
https://orcid.org/0000-0002-2750-5186
Sagui Celeste †
Roland Christopher *†
† Department of Physics, North Carolina State University, Raleigh, North Carolina 27695-8202, USA
‡ Department of Statistics, Florida State University, Tallahassee, Florida 32306, USA
* Email: cmroland@ncsu.edu.
08 09 2023
21 09 2023
08 09 2024
127 37 79077924
25 05 2023
11 07 2023
© 2023 The Authors. Published by American Chemical Society
2023
The Authors
https://creativecommons.org/licenses/by-nc-nd/4.0/ Permits non-commercial access and re-use, provided that author attribution and integrity are maintained; but does not permit creation of adaptations or other derivative works (https://creativecommons.org/licenses/by-nc-nd/4.0/).

Myotonic dystrophy type 1 is the most frequent form of muscular dystrophy in adults caused by an abnormal expansion of the CTG trinucleotide. Both the expanded DNA and the expanded CUG RNA transcript can fold into hairpins. Co-transcriptional formation of stable RNA·DNA hybrids can also enhance the instability of repeat tracts. We performed molecular dynamics simulations of homoduplexes associated with the disease, d(CTG)n and r(CUG)n, and their corresponding r(CAG)n:d(CTG)n and r(CUG)n:d(CAG)n hybrids that can form under bidirectional transcription and of non-pathological d(GTC)n and d(GUC)n homoduplexes. We characterized their conformations, stability, and dynamics and found that the U·U and T·T mismatches are dynamic, favoring anti–anti conformations inside the helical core, followed by anti–syn and syn–syn conformations. For DNA, the secondary minima in the non-expanding d(GTC)n helices are deeper, wider, and longer-lived than those in d(CTG)n, which constitutes another biophysical factor further differentiating the expanding and non-expanding sequences. The hybrid helices are closer to A-RNA, with the A-T and A-U pairs forming two stable Watson–Crick hydrogen bonds. The neutralizing ion distribution around the non-canonical pairs is also described.

National Institute of General Medical Sciences 10.13039/100000057 R01GM118508 document-id-old-9jp3c03538
document-id-new-14jp3c03538
ccc-price
==== Body
pmcIntroduction

Over 25 years ago, scientists realized that inherited neurological disorders known as “anticipation diseases,” where the age of the onset of the disease decreased and its severity increased, were caused by the intergenerational expansion of simple sequence repeats of 1 to 6 nucleotides.1−5 After a certain threshold in the length of the repeated sequence, the probability of further expansion and the severity of the disease increase with the length of the repeat. To date, approximately 50 DNA expandable diseases have been identified,6,7 of which the trinucleotide repeat (TR) diseases are the most common. The dynamic mutations in human genes with TR repeats can cause severe neurodegenerative and neuromuscular disorders known as trinucleotide (or triplet) repeat expansion diseases (TREDs).2,8−10 The expansion is believed to be primarily caused by some sort of slippage during DNA replication, repair, recombination, or transcription.5−7,11−15 Cell toxicity and death have been linked to the atypical conformation and functional changes of the transcripts and, when TRs are present in exons, of the translated proteins.6,16−25

In this work, we are interested in the structure of double helices formed by CTG expansions (CUG expansions in RNA). The CTG/CUG TRs lead to expansions associated with several diseases, while the GTC/GUC TRs do not exhibit pathological expansions. In particular, myotonic dystrophy is associated with an abnormal expansion of CTG (myotonic dystrophy type 1, DM1) and CCTG (myotonic dystrophy type 2, DM2). The disease is characterized by the progressive wasting and weakening of the muscles; patients with this disorder often have prolonged myotonia or muscle contractions which they are not able to relax after use.26,27 CTG repeats ranging between 5 and 38 are normal, while those ranging between 39 and 50 repeats are considered premutation alleles.28 Clinically affected individuals carry more than 50 repeats.29 The CTG TRs are located in the 3′-UTR of the dystrophia myotonica protein kinase gene while the CCTG repeats are found in the zinc finger 9 (ZNF9) gene.26,27,30 When transcribed, these form toxic RNA with CUG/CCUG (sense) and CAG/CAGG (antisense) repeats.31 The sense transcriptions fold into RNA hairpins which draw in cytoplasmic multiprotein complexes such as muscleblind-like 1 (MBNL1) which in turn cause muscle chloride channel dysfunction and abnormal insulin receptor regulation.30,32−35 The antisense transcriptions also represent a group of neurological disorders, including Huntington’s disease, and several kinds of spinocerebellar ataxias (SCAs).20,36 In addition, co-transcriptional R-loops (which consist of the hybrid RNA:DNA duplex formed by the template DNA and the complementary RNA strand, and the loose coding ssDNA) can cause further DNA damage and genome instability.37−39

In spite of the complexity of the molecular mechanisms behind TREDs, in all cases, stable, atypical DNA secondary structures have been identified as a “common and causative factor for expansion in human disease.”40 In the case of muscular dystrophy, the expanded repeats are located in non-coding regions, so that the disease is associated with toxic RNA gain-of-function. The first step toward understanding this disease therefore involves a structural and dynamical characterization of the atypical DNA and RNA structures (most commonly, hairpins), as well as hybrid duplexes that can form in an R-loop.

Experimental evidence demonstrate that r(CUG) and d(CTG) microsatellites can easily form hairpins.41−58 In particular, X-ray diffraction methods have been used to determine the crystal structures of RNA duplexes containing CUG repeats44,46,59,60 and both NMR and molecular dynamics (MD) were used to determine the structure of an RNA duplex with a single CUG repeat.45 Also, MD coupled to umbrella sampling was used to determine the structure and energetics of RNA duplexes containing CUG repeats.61 All these studies agree on the fact that CUG-containing RNA duplexes display the A-form,62,63 stabilized by the C·G and G·C Watson–Crick (WC) pairs (which form GpC steps when the CUG repeats are contiguous), whereas the mismatched U·U pairs are quite loose, resulting in a “stretched U·U wobble.”64 These mismatches are characterized by a hydrogen bond number ranging from zero to two, anti–anti conformations of the U·U pairs, and structures in which the U nucleotides are flipped out of the RNA helix altogether. Wobble structures have been observed in T·T pairs as well.65−67

In addition, CNG repeats are susceptible to form stable, long-lived R-loops due to the high thermal stability of rG-dC and rC-dG nucleotide pairs relative to dG–dC pairs.68,69 An R-loop is a three-stranded nucleic acid structure made up of a hybrid RNA:DNA duplex created by the template DNA and nascent RNA strand, as well as the displaced, non-template ssDNA. They were first detected during DNA replication and are also observed during transcription. R-loops have been extensively investigated because they can govern biological processes, including gene expression, DNA replication and repair, and immunoglobulin class-switch recombination.70,71 They may also damage DNA and create genetic instabilities, and they have recently been linked to neurological diseases, most likely by increasing the lifetime of single-stranded repeat DNA, allowing more non-B DNA secondary structures to develop and boosting repeat instability.37 Repeat tracts (CTG)n and (CAG)n, in particular, can be bidirectionally transcribed, allowing for single- and double-R-loop conformations in which either or both DNA strands can be RNA-bound.37−39 The determinants of trinucleotide R-loop formation are unknown, but the formation of stable RNA·DNA hybrids enhances the instability of CTG·CAG repeat tracts through aberrant processing or via slipped-DNA formation following RNA removal and its subsequent aberrant processing.39

Given the connection between neurodegenerative diseases and the associated secondary structures, the focus of this paper is to present a unified and comparative description of the structural and dynamical characteristics of the nucleic acid duplexes for both DNA and RNA based on CTG/CUG and GTC/GUC repeats, as well as hybrid double helices that can result from sense and antisense transcription of the expanding repeats, mainly d(CTG):r(CAG) and d(CAG):r(CUG). We note that our study does not address the transition from a single-stranded DNA or RNA to a hairpin structure. Indeed, one of the advantages of MD is that simulations can be made to start in any minimum of the free energy, independently of the complexity that took the molecule there. This study completes our previous efforts to present a comparative description of the nucleic acid duplexes that can form from TRs and other simple sequence repeats in all the possible reading frames. This work forms part our endeavor to characterize the atypical secondary structures of nucleic acids associated with TREDs. Previously, we focused on DNA and RNA homoduplexes, hybrid duplexes, triplexes, quadruplexes, Z-DNA, and hairpins associated with the most common TRs (CAG, CGG, CCG, GAA, and TTC) and the hexanucleotide repeats (GGGGCC, GGCCCC, and GGGCCT).72−82 In particular, we considered the DNA and RNA homoduplexes formed by CAG and GAC repeats72 complementary to the CNG and GNC (N=T or U) repeats in this study, and we harnessed the complementary role played by smFRET experiments and MD simulations in order to provide new insights into the role of sequence parity, TR interrupts, and favored type of loop structure on the dynamics of DNA CAG, GAC, CTG, and GTC hairpins.77,83 The present study made use of MD simulations complemented with the adaptively biased molecular dynamics (ABMD) method84 to calculate the free energy landscapes associated with U·U (T·T) mismatches for RNA (DNA). We note that strictly speaking, the non-canonical U·U pairs in RNA are not mismatches since RNA is not necessarily self-complementary. However, because we are considering both DNA and RNA in their helical form, we refer to these non-canonical basepair mismatches for ease of reference. The U·U mismatch conformation in RNA CUG has previously been explored theoretically61 and is included here to facilitate comparison with the other three mismatches (T·T in CTG and GTC, and U·U in GUC).

Methods

Molecular Dynamics Simulations

The pure RNA and DNA sequences investigated are shown in Figure 1a. We initially considered sequences with a single mismatch (5′-CCG-CUG-CCG-3′ and 5′-GGC-GUC-GGC-3′ for RNA termed r(CUG) and r(GUC) for short, and 5′-CCG-CTG-CCG-3′ and 5′-GGC-GTC-GGC-3′ for DNA termed d(CTG) and d(GTC) for short) in order to determine the most favorable U·U (RNA) and T·T (DNA) mismatch conformations, as well as their corresponding relative free energies via the ABMD method.84 In addition, we ran regular MD simulations for the sequences with various combinations of the χ angles for the mismatched sequences up to 1 μs. We extended these MD runs to the trinucleotide repeats 5′-(CUG)4-3′ and 5′-(GUC)4-3′ for RNA (termed r(CUG)4 and r(GUC)4 for short) and 5′-(CTG)4-3′ and 5′-(GTC)4-3′ for DNA (termed d(CTG)4 and d(GTC)4 for short) with various mismatch combinations. The length of these sequences represents a compromise between being short enough as to be computationally tractable and long enough as to capture the essential physics we are probing; i.e., the behavior of the stems of the hairpins in the presence of mismatches. As will be discussed later, duplexes of this length remain stable with few edge effects. We also note that hairpins may display long-time, large-scale dynamical behavior (e.g., strand slippage) that is beyond the scope of this study.77,83 Finally, stable A-RNA and B-DNA conformations from the above TRs were used to build r(CAG)4:d(CTG)4 and r(CUG)4:d(CAG)4 hybrids. The simulations were carried out using the PMEMD module of the AMBER v.1885 software package using the ff99 BSC186 for DNA ff99 BSC087 + χOL3 modification for RNA.88 The TIP3P model89 was used for the water molecules, along with the standard parameters for ions in the AMBER force fields.90 The long-range Coulomb interactions were evaluated by means of the Particle-Mesh Ewald (PME) method91 with a 9 Å cutoff and an Ewald coefficient of 0.30768. Similarly, the van der Waals interactions were calculated by means of a 9 Å atom-based non-bonded list, with a continuous correction applied to the long-range part of the interaction. The production runs were generated using the leap-frog algorithms with a 2 fs timestep in an NPT ensemble. A Langevin thermostat with a collision frequency of 1 ps–1 and Monte Carlo barostat was used to maintain the pressure of the system to 1 atm. The SHAKE algorithm was applied to all bonds involving hydrogen atoms. The MD simulations were 1 μs long and typically involved different initial values for the χ glycosyl torsion angles, with conformations saved every picosecond. These conformations were then analyzed for their dynamical and structural characteristics (e.g., handedness as defined via Figure S1) as discussed in the subsequent sections.

Figure 1 (a) Schematic of sequences considered in this study (for both DNA and RNA): r(CUG), r(CUG)4, d(CTG), and d(CTG)4. Structures for r(GUC) and d(GTC) are similar and may be obtained by interchanging G and C in the illustrated sequences. (b) χ angle of the U base is the O4′-C1′-N1-C2 dihedral (similar for T in DNA). (c) Schematic view of the center-of-mass pseudodihedral angle Ω (for U14 in r(CUG)).

Free Energy Maps

The sequences with a single mismatch, r(CUG), d(CTG), r(GUC), and d(GTC), were used to identify the mismatch conformation that minimizes the free energy. To calculate the free energy maps, we made use of the ABMD method84,92 as implemented for PMEMD in AMBER v.18.85 ABMD is a non-equilibrium MD method that belongs to the general category of umbrella sampling methods with a history-dependent biasing potential, that in the long-time limit reproduces the negative of free energy. The free energy – or potential of mean force (PMF) – is calculated as a function of one or more collective variables, which must carefully be chosen as to reflect the underlying physics of the problem. The free energy maps of these mismatches were calculated as a function of two main collective variables carefully chosen to reflect possible mismatch conformations and to compare with previous TR investigations.72,75 Specifically, we considered collective variables based on the following torsion angles: (1) Ω14 as the center-of-mass (COM) pseudodihedral angle, which is defined using the COMs of four atom groups: G6(C1′, C2′, C3′, C4′, O4′), C13(C1′, C2′, C3′, C4′, O4′), G15(C1′, C2′, C3′, C4′, O4′), and U14(N1, C2, O2, N3, C4, O4, C5, C6) for RNA, and similarly for DNA using T14 instead of U14. Likewise, Ω5 is similarly defined for U5 for RNA and T5 for DNA. This variable describes the base unstacking of U or T with respect to the helical axis; (2) χ5 as the glycosyl torsion angle χ of U5 or T5, namely the dihedral angle O4′–C1′–N1–C2; and (3) χ14 which represents the χ angle of U14 or T14. These collective variables probe the mismatch conformations inside the helical core. A schematic view of these collective variables is shown in Figure 1b.

With these variables, we constructed three free energy landscapes, (Ω14, χ14), (χ5, χ14), and (Ω5, Ω14). For the first combination, (Ω14, χ14), we found that if we choose χ5 in the anti range, U5 (or T5) stays in its anti-conformation for all calculations, so this free energy map explores anti–anti and anti–syn conformations of T5·T14 and U5·U14, and whether they remain inside the helical core or not. By construction, this diagram cannot explore syn–syn conformations. The (χ5, χ14) diagram, on the other hand, can explore all options of χ (anti–anti, anti–syn, syn–anti, and syn–syn) but is degenerate with respect to Ω, i.e., it cannot tell whether the bases are inside the helix or have flipped out. Finally, the (Ω5, Ω14) combination tests whether both bases flip out. After the initial conformations were set up as explained below, multiple-walker ABMD runs with eight walkers at 300 K in the NVT ensemble were employed. A sequence of three different runs was used in order to successively refine the free energy maps. First, a 200 ns WT-ABMD simulation with parameters τF = 1 ps, 4Δξ = 0.5 radians, and pseudo-temperature 10,000 K provided for a rough picture of the map. This was followed by an intermediate resolution run of 300 ns duration (parameters τF= 1 ps, 4Δξ = 0.2 radians, and pseudo-temperature 10,000). The final production run lasted at least 200 ns with parameters similar to the intermediate runs but with a flooding time scale of τF = 5 ps. For these runs, the total number of hydrogen bonds in neighboring C–G WC base pairs as well as root-mean-squared deviation of phosphate atoms of C–G W–C base pair were restrained in order to avoid the large-scale twisting of the whole structure during the long simulations. Additionally, the distance between center of mass of adjacent bases was slightly constrained to provide the required room for the rotation of the U and T bases. These constraints, however, were chosen to be flexible enough so as to readily allow for the probing of the single mismatch conformations. A given free energy landscape was deemed to have converged when both the position and differences in the free energy values of the minima remain approximately constant as further ABMD cycles were performed. For RNA (DNA), about 700 ns (1500 ns) are required for each of the (χ5, χ14) maps; the DNA landscapes are harder to converge because of the greater flexibility of DNA.

Initial Conformations

Initial conformations for one- and four-repeat sequences were created as follows. First, we created the duplexes with the four possible combinations of χ angle for the mismatches: anti–anti, anti–syn, syn–anti, and syn–syn. These were then solvated in an octahedral box with an appropriate number of neutralizing Na+ ions as well as 0.15 M salt as in previous work,72,75,93 with a distance of at least 10 Å between the duplexes and walls of the box. The box was then filled with a suitable number of waters. The system was then minimized: first keeping the nucleic acid and ions fixed; then, allowing them to move. Subsequently, the temperature was gradually raised using constant volume simulations from 0 to 300 K over 50 ps, followed by a further 50 ps run. Then a 100 ps run at constant volume was used to gradually reduce the restraining harmonic constants for nucleic acids and ions. This was followed by a 1.0 ns constant pressure run, with the χ angles of U5 and U14 (or T5 and T14) slightly restrained so that these retain their initial anti- or syn-conformation. We took random conformations from the last 100 ps of these runs as the initial conformations for both the ABMD and MD runs. In particular, for the (Ω14, χ14) phase diagrams (where the collective variables are angles associated with U14 or T14), we picked four structures from U5(anti)-U14(anti) (or T5(anti)-T14(anti)) and four from U5(anti)-U14(syn) (or T5(anti)-T14(syn)) since the aim is to assess the anti–syn flipping of U14 (or T14), which is completely equivalent to U5 (or T5). Since DNA is less stable than RNA, a small restraint was applied to χ5 for the (Ω14, χ14) DNA diagram. For the (χ5, χ14) diagrams, we picked two structures from each of the four runs (anti–anti, anti–syn, syn–anti, and syn–syn). For the four repeats, r(CUG)4, d(CTG)4, r(GUC)4, and d(GTC)4, we considered four conformations anti–anti, anti–syn, syn–anti, and syn–syn with mismatches all inside the helical core, followed the same minimization and equilibrium steps described above and then 1 μs simulations were run at 303 K. The r(CAG)4:d(CTG)4 and r(CUG)4:d(CAG)4 hybrids were each built in ideal A-RNA or B-DNA conformation so that we had four initial duplexes. After equilibration, a total of 1 μs of MD was run at 303 K for each hybrid duplex.

Results

Free Energy Maps for Single Mismatches

We begin our discussion with a consideration of a single-mismatch in the CUG/GUC RNA and CTG/GTC DNA sequences shown in Figure 1a. For each duplex, we computed three free energy landscapes using the collective variables based on the Ω and χ torsion angles employed in previous studies;61,72 the (Ω14, χ14) landscape is “asymmetric”, while the (χ5, χ14) and (Ω5, Ω14) maps are symmetric (when the calculations have converged). Values of Ω around −45 to 100 degrees represent bases inside the helix core, while angles beyond these values indicate bases that have flipped out. Positive (negative) values here correspond to a flipping to the major (minor) groove direction, respectively. Values of χ between – 180° and – 90° (likewise between 90° and 180°) are considered anti conformation, while the other half range – 90° to 90° corresponds to syn conformations. These free energy landscapes are characterized by several stable minima. We have set the deepest minimum value to zero and have marked the most prominent ones with letters. The location and values of these minima are summarized in Table 1, while conformations corresponding to these minima are given in Figure 2a–c for U-mismatches and Figure 3a–c for T-mismatches.

Figure 2 Different conformations of U·U mismatches in the local minima of r(CUG) free energy landscape. Here, A conformations are all anti–anti, while B conformations represent anti–syn and C is the syn–syn conformation. The H-bonds are also shown.

Figure 3 Different conformations of T·T mismatches in the local minima of d(CTG) free energy landscape. Here, A represents the anti–anti conformation, B represents the anti–syn, and C represents the syn–syn conformations. Observed H-bonds are shown.

Table 1 Main Minima for All the Free Energy Mapsa

mismatch form	anti–anti	anti–syn	syn–syn	
main H-bond	single: O4-N3:H3, double: N3:H3-O2,O4-N3:H3	single: N3:H3-O4	–	
r(CUG)	approximate location (Ω14, χ14)	(68,200) for A1, (16,194) for A2	(33,50) for B1, (−50,65) for B2	–	
relative free energy (kcal/mol)	0 for A1, 0.3±0.1 for A2	6.1±0.2 for B1, 7.9±0.2 for B2	–	
approximate location (χ5, χ14)	(−160,–162)	(−157,49) for B, (49,–154) for B′	(58,58)	
relative free energy (kcal/mol)	0	5.8±0.2 for B, 5.8±0.1 for B′	11.2±0.2	
approximate location (Ω5, Ω14)	(6,49) for A1, (49,9) for A1′, (70,70) for A2	–	–	
relative free energy (kcal/mol)	0 for A1, 0.1±0.4 for A1′, 0.2±0.4 for A2	–	–	
 	
mismatch form	anti–anti	anti–syn	syn–syn	
main H-bond	double: N3:H3-O2,O4-N3:H3	single: N3:H3-O4	–	
d(CTG)	approximate location (Ω14, χ14)	(70,226)	(39,68) for B1, (108,65) for B2	–	
relative free energy (kcal/mol)	0	7.0±0.2 for B1, 8.4±0.1 for B2	–	
approximate location (χ5, χ14)	(−124,–110)	(−122,67) for B, (67,–119) for B′	(69,64)	
relative free energy (kcal/mol)	0	6.4±0.6 for B, 7.1±0.2 for B′	12.3±0.5	
approximate location (Ω5, Ω14)	(49,76) for A, (77,51) for A′	–	–	
relative free energy (kcal/mol)	0 for A, 0±0.2 for A′	–	–	
 	
mismatch form	anti–anti	anti–syn	syn–syn	
main H-bond	single: O4-N3:H3, double: N3:H3-O2,O4-N3:H3	single: N3:H3-O4	–	
r(GUC)	approximate location (Ω14, χ14)	(64,205) for A1, (27,202) for A2	(42,53) for B1, (−56,67) for B2	–	
relative free energy (kcal/mol)	0 for A1, 0.3±0.4 for A2	5.0±0.5 for B1, 7.7±1.1 for B2	–	
approximate location (χ5, χ14)	(−158,–158)	(53,–154) for B, (−158,71) for B′	(56,67)	
relative free energy (kcal/mol)	0	5.5±0.2 for B, 6.0±0.7 for B′	11.1±0.2	
approximate location (Ω5, Ω14)	(27,64) for A1, (64,35) for A1′, (71,71) for A2	–	–	
relative free energy (kcal/mol)	0 for A1, 0.2±0.2 for A1′, 1.0±0.3 for A2	–	–	
 	
mismatch form	anti–anti	anti–syn	syn–syn	
main H-bond	double: N3:H3-O2,O4-N3:H3	single: N3:H3-O4	–	
d(GTC)	approximate location (Ω14, χ14)	(39,238)	(44,56) for B1, (99,61) for B2	–	
relative free energy (kcal/mol)	0	1.82±0.92 for B1, 4.05±0.83 for B2	–	
approximate location (χ5, χ14)	(−136,–136)	(−136,49) for B, (35,–140) for B′	(49,49)	
relative free energy (kcal/mol)	0	3.4±0.2 for B, 3.5±0.3 for B′	7.1±0.4	
approximate location (Ω5, Ω14)	(45,85) for A, (82,45) for A′	–	–	
relative free energy (kcal/mol)	0 for A, 0.6±0.1 for A′	–	–	
a The primed letters represent the mirror images of the corresponding unprimed letters. All the values and errors are calculated based on the last 100 ns of the ABMD simulations. Positions of minima are given in degrees, and the free energy is kcal/mol. Note that the ABMD free energy differences and errors were calculated as the average and standard deviation of the sample of multiple PMFs obtained from regular intervals of the last 100 ns of the simulations where we estimate that systems have explored all the relevant states.

Figure 4 gives the calculated free energy maps for r(CUG) (Figure 4a–c) and d(CTG) (Figure 4d–f), and Figure 5 does the same for r(GUC) (Figure 5a–c) and d(GTC) (Figure 5d–f). The landscapes resemble each other and share common features. Consider the (Ω14, χ14) landscapes. The deeper minima are associated with – 10° < Ω < 90° which correspond to well-stacked bases inside the helical core; minima associated with other Ωs correspond to bases that have flipped out, and these are considerably shallower. In all cases, the deepest minima A1 for RNA and A for DNA correspond to anti–anti conformations with zero to two hydrogen bonds as shown in Figures 2a and 3a. We note that the anti–anti minimum on the (Ω14, χ14) landscape of r(CUG) is somewhat broader than the corresponding d(CTG) anti–anti valley. This is due to the presence of additional r(CUG) anti–anti conformations shown in Figure 2, which are characterized by two, one, and no hydrogen bonds. Another feature is that for the RNA duplexes, the anti–anti valley is located primarily in the 180° ≤ χ14 ≤ 205° range, which is proper anti–anti, while for d(CTG), it is 220° ≤ χ14 ≤ 270° which is associated with a high anti range. Such a feature was also noted in a recent study on CAG and GAC TR structures.72 The lower values in RNA can be explained by the presence of the additional hydroxyl group at the 2′ position in the sugar ring of RNA (structure shown in ref (72)). The effect of this hydroxyl group is to interact with the RNA backbone pulling the sugar ring at one end and giving rise to a twist at the other. All in all, this results in an overall decrease in the χ-angle associated with r(CUG) and r(GUC). Transitions between syn–anti (B) and anti–anti (A1,A2) take place either by a direct change in the χ angle with the nucleotide remaining inside the helical core or by extruding the nucleotide toward either the major or minor groove and then flipping back toward the inner core with a changed χ angle. Similar results are also observed for the corresponding r(GUC) and d(GTC) duplexes shown in Figure 5a,d.

Figure 4 Free energy maps for single mismatches in r(CUG)(a–c) and d(CTG)(d–f) with different choices of collective variables. (a,d):(Ω14, χ14); (b,e):(χ5, χ14); and (c,f):(Ω5, Ω14). The values in the bar are given in kcal/mol. Each contour line approximately shows 1 kcal/mol.

Figure 5 Free energy maps for single mismatches in r(GUC)(a–c) and d(GTC)(d–f) with different choices of collective variables. (a,d):(Ω14, χ14); (b,e):(χ5, χ14); and (c,f):(Ω5, Ω14). The values in the bar are given in kcal/mol. Each contour line approximately shows 1 kcal/mol.

The (χ5, χ14) free energy maps are shown in Figures 4b,e and 5b,e. As expected, these maps display mirror symmetry across the diagonal. Primed letters indicate minima related by mirror symmetry (e.g., B denotes an anti–syn conformation and B′ the corresponding syn–anti conformation). We have identified two anti–syn minima B1 and B2 on the (Ω14, χ14) map, with structures as shown in Figures 2b and 3b. On this map that only has one χ, these conformations are degenerated. Examining the (χ5, χ14) free energy landscapes, we readily see that the anti–anti conformation corresponds to the deepest minima, followed by anti–syn and then the syn–syn conformation. The fact that d(CTG) and d(GTC) are associated with (high-anti)–(high-anti) conformations in contrast to the anti–anti conformations in r(CUG) and r(GUC) is visually evident here. These features are due to the greater flexibility of the DNA sugar ring, which allows the T·T mismatches to explore a somewhat larger range of conformations. Unfortunately, this feature also slows down the numerical convergence of the DNA free energy maps.

To investigate the correlation between the stacking and unstacking of the mismatches, we calculated the (Ω5, Ω14) free energy landscapes shown in Figures 4c,f and 5c,f. As expected, these free energy maps show mirror symmetry. There are two broad, light-green channels that cross each other approximately located at 0° ≤ Ω ≤ 90° for both Ω14 and Ω5. These light-green regions (but not the purple ones) indicate that unstacking of each uridine or thymidine is anti correlated: when one uridine or thymidine begins to unstack, the other will stay put inside the helical core, and vice versa. The minima in the purple valleys correspond to mismatches inside the helical core in anti–anti conformations. The minima A1 and A1′ (which are characterized by either two or one hydrogen bonds, Figure 2a) are well away from the A2 minima. The coexisting conformations with zero, one, and two hydrogen bonds are indistinguishable within limits of ABMD sampling; however, they are separated by a small but distinct free energy barrier running along the diagonal separating the unprimed and primed conformations. Using the reweighting technique,94 we estimate that the single hydrogen-bond conformations have the lowest free energy, while zero and double hydrogen bonds are higher around 0.2 and 0.9 kcal/mol, respectively, with an average error of 0.1 kcal/mol estimated from the block-analysis method.94,95 It is important to note that although certain reported free energy differences in hydrogen bond free energy fall below the thermal energy, they are in line with previous reports (including an umbrella sampling investigation of r(CUG) structure)61,96,97 that zero and single hydrogen bonds are preferred to double hydrogen bonds, most likely due to the smaller size of U-U pairs compared to WC pairs. Moving to the d(CTG)/d(GTC) maps, these display similar features as r(CUG)/r(GUC) but with shorter valleys for the anti–anti conformations, with the absolute minima displaced toward higher values of Ω. This is consistent with only a single anti–anti conformation conformation for the T·T mismatch with two hydrogen bonds as shown in Figure 3a. Other anti–anti conformations, say with a single hydrogen bonds, are simply too close to preclude an unambiguous identification. The most important difference between the CNG and GNC maps occurs in DNA (Figures 4e and 5e), where the anti–syn, syn–anti, and even syn–syn relative minima in d(GTC) are considerably deeper and wider than those in d(CTG).

To gain further insight into the behavior of the double helices with the mismatches, we carried out a single, regular 1 μ MD simulation run for each of the homoduplexes (including homoduplexes with one and four mismatches) and different initial conditions, as shown schematically in Figure 1a. The U·U/T·T mismatches were initialized as being all inside the helical core with a conformation that is either anti–anti, anti–syn, or syn–syn; i.e., one separate run for each of these initial conditions. For single mismatch homoduplexes, we observed that mismatches in an anti–anti conformation remain as such throughout the duration of the simulation, reflecting the fact that it is the lowest free energy conformation. For other conformations, we recognized transitions to anti–anti conformations within the simulation time frame. For example, top panel SI Figure S2b–d shows such transitions for the r(CUG) mismatches started in the syn–syn and syn–anti conformations. While not all simulated structures were observed to transition to their anti–anti conformation, we ultimately expect all to end up there given enough simulation time. Bottom panel for Figure S2a–d also shows time traces of the hydrogen bond number for the different mismatches: these vary from two to zero reflecting the different conformations shown in Figure 2a–c and provide for a visual appreciation of their relative population. The plots of glycosidic torsion angle χ and hydrogen bond number as function of time for some of the other single mismatched double helices are shown in top and bottom panels for SI Figure S3a–d.

We now turn to the mismatch behavior in the TRs r(CUG)4, r(GUC)4, d(CTG)4, and d(GTC)4. Again, we ran a single simulation for each duplex with the dihedral χ angles of each mismatch in either an initial anti–anti, anti–syn, or syn–syn conformation. Figure S4 shows RMSD of all atoms of these helices with respect to the average frame of the last 200 ns as a function of time; while Figure 6a–f and SI Figures S5a–f, S6a–f, and S7a–f give time traces of the χ angle and number of hydrogen bonds. For all these double helices, the anti–anti conformations were found to be the most stable, as expected. The simulations of all the other structures showed evidence of transitions to the anti–anti state to varying degrees. Here, we give a qualitative summary of the observed transitions. First, Figure S4 shows the RMSD of all atoms in the helices with respect to the average frame of the last 200 ns as a function of time. The r(CUG)4 anti–anti exhibits relatively larger local and global conformational fluctuations. The main source of fluctuations was related to transitions between two, one, and no hydrogen bond conformations; Figure 6a–f shows time traces for the χ angles and number of hydrogen bonds, with similar plots for the other structures being relegated to SI Figures S5a–f, S6a–f, and S7a–f. While all duplexes initialized in their anti–anti conformation remained stable, transitions were observed for structures initially in their anti–syn and syn–syn conformations. For r(CUG)4 in initial syn–syn conformation, three base pairs transitioned to anti–anti and one moved to anti–syn after approximately 500 ns. The initial anti–syn conformation showed some instability, with some base pairs, including WC ones, losing hydrogen bonding and stacking. For r(GUC)4 in an initial syn–syn conformation, we observed three mismatches transitioning to anti–syn and one to anti–anti after around 300 ns; in contrast, the helix in the initial anti–syn conformation quickly lost its helical form and displayed some unwinding and stretching, ultimately transitioning to anti–anti. For d(CTG)4 in an initial syn–syn conformation, we observed two mismatches transitioning to anti–anti in under 40 ns, and two others to anti–syn after around 400 ns, with some mismatches opening up. Likewise for d(CTG)4 in initial anti–syn, three mismatches transitioned to anti–anti (two in under 50 ns and one after 750 ns). Similarly, for d(GTC)4 initially in syn–syn, we noticed three transitions to anti–syn, causing a sharp helical bend, but no χ transitions were observed for d(GTC)4 helices initially in their anti–syn conformation. In terms of hydrogen bonds analysis, the results were found to be qualitatively similar to the helices with a single mismatch. For both RNA helices in the initial anti–anti conformation, the system primarily wobbled between the conformations with two, one, and no hydrogen bonds, shown in Figure 2a, while the DNA structures, particularly d(CTG)4, formed two relatively stable hydrogen bond structures resulting in a shorter T·T mismatch distance which in turn is reflected in the RMSD shown in Figure S4a.

Figure 6 Results from 1 μs MD simulation for r(CUG)4 helices. Shown are angle χ (left) and number of hydrogen bonds (right) associated with the internal mismatches of the helices. Initial conformations for individual panels (top to bottom) are (a) anti–anti; (b) anti–syn; and (c) syn–syn. Each individual panel, (a), (b), or (c) is divided in two rows. The upper row corresponds to U5·U20 while the lower row corresponds to U8·U17. χ5 and χ8 are black, and χ20 and χ17 are red. On the right, different colors represent different numbers of hydrogen bonds (0, 1, or 2).

Structural Characteristics, Dynamical Fluctuations, and the Principal Component Analysis (PCA)

We now turn to the dynamical fluctuations of the different helices, which were analyzed in terms of the PCA. We initially focus on the single mismatch systems as these contain the key dynamical features inherent in the multi-mismatch helices. PCA results for the U·U (T·T) mismatch of r(CUG) (r(GUC)) and d(CTG) (d(GTC)) systems are shown in Figures 7a,b (SI Figure S8a,b) and 8a,b (SI Figure S9a,b). The results are based on the 1μs MD runs with the mismatches in the anti–anti conformation. We found that the contribution of the first two eigenvectors correspond to the mismatch base fluctuations and account for more than 79% (r(CUG)), 45% (d(CTG)), and 78% (both r(GUC) and d(GTC)); projections of these eigenvectors for r(CUG) and d(CTG) are shown in the figures. For r(CUG), the projection of the first eigenvector is rather broad and flat without any obvious peaks. A check of the trajectories verifies that these fluctuations are mostly due to a shearing and opening of the U·U mismatch. By comparison, the projection of the second eigenvector (also describing a shearing and opening of the mismatch) displays two distinct peaks on either the positive and negative side. As shown in Figure 7b, the two peaks represent the one and two hydrogen bonds and their mirror image conformations in Figure 2a. Two subpeaks can also be observed on the positive side that represent fluctuations between the same conformations; these two subpeaks seem to coalesce together on the negative side. For d(CTG), the projection on the first eigenvector gives two peaks, which represents the fluctuations of sugar ring of T5 base (Figure 8b). This results in large fluctuations of the δ, ϵ, and ζ backbone torsion angles of T5, as well as a fluctuation of the sugar pucker which oscillates between C2′-endo and C3′-endo. Similarly, the projection of displacements along the first and second eigenvector for r(GUC) are bimodal and unimodal, respectively (SI Figure S8a,b). The peaks in the first eigenvector represent the single and two hydrogen bond conformations where the Ω angle (opening) changes in two different directions, while the single peak associated with the second eigenvector plot corresponds to the mismatch opening toward the same direction. For d(GTC), the histograms of the projections of displacements along the first and second eigenvectors are similar to those of d(CTG), with the presence of two peaks along the first eigenvector (see SI Figure S9a). However, in d(GTC), the two peaks and their distance along the first eigenvector are more pronounced, indicating that the T·T bases spend more time in either of the peaks. In addition to torsion angle fluctuations, we observe base shearing and opening which result in fluctuations between one and two hydrogen bonds.

Figure 7 Results of the PCA analysis for the U·U mismatch in anti–anti conformation for r(CUG). (a) Histogram of projections on the first (cyan) and second (grey) eigenvectors. (b) Fluctuations of conformation along the second eigenvector. Cyan and red structures highlight the A1 and A1′ (mirror image of A1) structures as in Figure 2.

Figure 8 Results of PCA analysis for the T·T mismatch in anti–anti conformation for d(CTG). (a) Histogram of projections on the first (cyan) and second (grey) eigenvectors. (b) Fluctuations of conformation along the first eigenvector. Cyan and red structures show the two conformations with the largest RMSD.

Having understood the PCA behavior of single mismatch systems, we turn to the double helices with multiple mismatches. We applied the PCA method to the backbone atoms of r(CUG)4, r(GUC)4, d(CTG)4, and d(GTC)4 helices in their initial anti–anti conformations with results averaged over the last 200 ns. Our results show that the dynamical fluctuations associated with the r(CUG)4 double helix is slightly simpler in nature, insofar as the first three PC account for more than 80% of the variance; by contrast, five PCs are required for the same level of variance for d(CTG)4. The dynamical fluctuations of r(GUC)4 and d(GTC)4 appear to be rather similar with the first three PCs accounting for 78% of the fluctuations. A visual examination of the trajectories created by the PCA shows that the primary essential movements correspond to a bending/unbending and winding/unwinding of the helices, which are related to changes in the roll, tilt, and twist, as these motions are correlated, as well as an inclination of the mismatched steps. Figure 9 illustrates the aforementioned conformational fluctuations along the direction of the first PCA eigenvector. Scatter plots of PC1 and PC2 for the duplexes shown in SI Figure S10 indicate that both RNA helices, especially r(CUG)4, undergo more pronounced motions (sample more regions) as compared to the DNA helices, which is also in agreement with the RMSD results.

Figure 9 Conformational fluctuations around the first eigenvector direction based on the PCA analysis of the backbone of (a) r(CUG)4, (b) d(CTG)4, (c) r(GUC)4, and (d) d(GTC)4.

The conformational changes undergone by the mismatch duplexes can also be quantified using the notion of handedness (defined in the SI) and the radius of gyration. By definition (see Figure S1), handedness is positive for right-handed helices such as B-DNA and A-RNA, zero for a straight duplex, and negative for left-handed helices such as Z-DNA. As such, it is closely related to the “helical twist” (which combines the structural parameters of twist and bending angle) when shift and slide are negligible for a particular step. The radius of gyration characterizes the compactness of the helix and is comparable to the “helical rise” of a step. These quantities are given in the SI: SI Figures S11 and S12; we see that the handedness for both r(CUG)4 and d(CTG)4 are somewhat reduced so that these helices are somewhat unwound when compared to duplexes without mismatches. We note that the population peak associated with r(CUG)4 is quite broad when compared to that of the other helices, again reflecting the associated PCA results. Similar results are obtained for radius of gyration plots (bottom, SI Figure S11). In contrast, the d(CTG)4 radius of gyration remains relatively constant throughout the duration of the simulation. The results for r(GUC)4 and d(GTC)4 are relatively similar and shown in SI Figure S12.

To characterize the structural aspects of the mismatched helices, we used the program 3DNA98 with data taken from the last 200 ns of the simulations; only the mismatched bases and middle seven base-pair steps were considered. Results are given in Table 2 and Figure 10a,b. Generally speaking, short nucleic acid strands are not as rigid as infinitely long ones and the appearance of the mismatches causes significant distortions.99 Our results suggest that among different structural parameters, shear, stretch, stagger, as well as opening vary the most when matched base pairs A-U and A-T are replaced by mismatches U·U and T·T, which is in agreement with the PCA analysis. The base-pair step parameters shown in Figure 10 roughly show the step periodicity although a few results do not properly reflect the symmetry expected for the sequence due to the fact that the curves are averaged over only the last 200ns of the simulations, and there are long-lived conformational metastable states in the dynamic behavior of the helices. In general, we see that twist and helical rise seem to be distorted the most and are correlated in r(CUG)4. The average twist for r(CUG)4 drops to 27.1°, while the average helical rise increases to 3.4 Å. These changes, combined with an increase in the major groove width (21.4 Å), are the characteristics of an unwound and stretched helical structure. The r(GUC)4 exhibits similar, but less pronounced, behavior. On the other hand, d(CTG)4 and d(GTC)4 are closer to that of a standard B-DNA helix, except that some of the mismatched steps are a bit distorted. Roll and twist have been found to be inversely correlated.100,101 Furthermore, tilt primarily fluctuates with values close to zero. These two parameters may conveniently be combined to form the local bend angle () as a probe of the bending and unbending of the duplexes. According to this parameter, both d(CUG)4 and d(CTG)4 are slightly less bent than A-RNA and B-DNA without mismatches. The average total bending angle (in degrees) of the structures was obtained using the curvilinear helical axis from the software Curves+102 as follow: d(CUG)4: 21.5±10.3, d(CTG)4: 15.6±8.2, r(GUC)4: 18.1±9.3, and d(GTC)4: 15.6±8.4. Thus, both RNA helices are characterized by a slight bending and stretching. We can expect the bending angle to increase as the length of the structures increases.97 We note that the base-pair steps with mismatches systematically have lower inclination in comparison to A-RNA and B-DNA without mismatches. The glycosidic torsion angle χ remained in the anti–anti range throughout of simulation, with average values close to the free energy results. The WC C–G base pairs also generally fell into the anti range, except for some terminal bases which tended to display large fluctuations.

Figure 10 Comparison of step parameters for homoduplexes and WC helices. Average twist, inclination, bend (), rise, slide, and Zp for middle base-pair steps of homoduplexes r(CUG)4, r(GUC)4, d(CTG)4, and d(GTC)4 (all with mismatches in anti–anti conformation); and of standard helices r(CUG:CAG)4, r(GUC:GAC)4, d(CTG:CAG)4, and d(GTC:GAC)4. Colors are used as follows. (a) Red: d(CTG:CAG)4; orange: r(CUG:CAG)4; blue: d(CTG)4; green: r(CUG)4; and (b) red: d(GTC:GAC)4; orange: r(GUC:GAC)4; blue: d(GTC)4; green: r(GUC)4. Data were averaged over the last 200 ns. Here, ‘X’ in the horizontal label indicates either T, U, or A, depending on the sequence being considered.

Table 2 Average Structural Parameters of Mismatchesa

 	opening	net non-planarity	shear	stretch	stagger	χ	
U·U (r(CUG)4)	5.38	14.55	1.71	0.55	0.14	–158.2	
T·T (d(CTG)4)	9.35	13.2	2.86	0.96	0.46	–120.05	
U·U (r(GUC)4)	5.21	13.97	1.39	1.06	0.18	–155.96	
T·T (d(GTC)4)	4.40	12.37	2.43	1.07	0.24	–124.2	
A-U (r(CUG:CAG)4)	0.92	13.1	0.10	0.02	0.06	 	
A-T (d(CTG:CAG)4)	0.86	10.92	0.01	0.07	0.09	 	
a Average base-pair parameters as obtained from the last 200 ns of the simulations. All the angles, i.e., opening, net non-planarity, and χ, are given in degrees and all displacements, i.e., shear, stretch, and stagger are given in terms of Angstroms. The net non-planarity angle is defined as square root of the buckle squared plus propeller squared. As shear, stretch, and stagger accept both positive and negative values, their root mean square value was used.

Distribution of Neutralizing Ions Around the Mismatched Helices

An important structural aspect of nucleic acids is their polyanionic nature such that water and counterions are crucial for their stability. In solution, counterions surround the nucleic acid structure and neutralize the helices’ anionic phosphates. They can also establish water mediated contacts and also, albeit less frequently, direct contacts with the electronegative groups. Hence, the counterion distribution is a key part of the structure and stability of the nucleic acid conformations. We have thus characterized the neutralizing Na+ ion distribution around the mismatched helices, with a focus on identifying the so-called binding sites associated with the mismatches. Figures S13a–d (top and bottom panels) and S14a–d (top and bottom panels) show the distance between the Na+ ions to the center of mass of the single mismatches. Here, the different colors represent different ions in order to visually bring out the single-ion binding time. Ions within a distance of 5 Å always have a direct interaction with the base mismatches. Visual observations show that for r(CUG) and r(GUC), the time an ion spends around a U·U mismatch are roughly similar for the different conformations, except for the syn–syn conformation in r(GUC), where the time appears to be somewhat longer. A similar observation was previously found to be a characteristic of syn–syn C–C mismatches for B-DNA GCC structures.75 Likewise, the occupation times for T·T mismatches in d(CTC) and d(GTC) appear to be very similar and slightly longer than those for the U·U mismatches.

Figures 11 and 12 show the ion occupancy for a single U·U mismatch for RNA in r(CUG)4 (left panel, Figure 11a–c) and r(GUC)4 (right panel, Figure 11a–c) and T·T mismatch for DNA in d(CTG)4 (left panel, Figure 12a–c) and d(GTC)4 (right panel, Figure 12a–c). Considering the U·U mismatches in an anti–anti conformation, the ions are primarily associated with the major groove O4 atoms (atomic structures depicted in Figure S15a,b); this binding site may also involve the O6 atom of the neighboring G base and is illustrated in Figure 13a. This ionic interaction does not disturb the base/sugar orientation of the uridines, so that these always remain in their anti–anti conformations. In addition, there are considerably smaller ion occupancies associated with O2 and the backbone phosphates OP1 and OP2. For U·U mismatches in the anti–syn (or syn–anti) and syn–syn conformations, we observe substantial O4 binding but also considerably more at the O2 site. Roughly speaking, there is inversion symmetry between the contributions from the U bases in the anti–syn and syn–anti conformations. The two binding sites in major groove of the anti–syn conformation have also been previously reported.61 This type of binding appears to slow down the transition to the equilibrium glycosidic conformation. Since the syn–syn is not a stable conformation, the high occupancy binding is ultimately transient. Figure 11 compares the binding around U·U mismatches for CUG and GUC helices, which are relatively similar. Finally, SI Figure S16a,b gives a visual presentation of the most prominent ion density around the U·U mismatches.

Figure 11 (a–c) Average ion occupancy around U·U mismatches in r(CUG)4 (left) and r(GUC)4 (right). Blue: bases in the first strand. Red: bases in the second strand. Different panels are for different initial conformations of the mismatches. The averages are based on data taken from the last 200 ns of the simulations.

Figure 12 (a–c) Average ion occupancy around T·T mismatches in d(CTG)4 (left) and d(GTC)4 (right). Blue: bases in the first strand. Red: bases in the second strand. Different panels are for different initial conformations of the mismatches. The averages are based on data taken from the last 200 ns of the simulations.

Figure 13 Illustration of two main Na+ ion binding sites for U·U and T·T mismatches. U·U/T·T mismatches are highlighted in cyan color, and Na+ ions are represented by orange spheres. (a) For r(CUG) in anti–anti conformation, there is a binding site in the major groove, where the ion binds to the O4 atoms in each U base of the mismatch, and to the neighboring G:O6 atom. (b) For d(CTG) in anti–anti conformation, there is a binding site in the minor groove where Na+ binds to T14:O2 and neighboring G15:N3.

Turning to d(CTG) helices, Figure 12 shows that the most prominent ion binding site is associated with O2, whose atomic conformation is shown in Figure 13b. This is true for all combinations of initial conditions for the mismatches. As with the U·U mismatches, the T·T mismatches also show contributions from the backbone OP1 and OP2 sites. Comparing the CTG and GTC results, we see that while the binding sites for the anti–anti conformations roughly correspond, the O4 contributions for the anti–syn and syn–syn for GTC is more substantial than that of CTG. This is most probably due to the different steps surrounding the mismatches. A visual representation of the corresponding ion densities is shown in SI Figure S17a,b.

MD Investigations of Hybrid Duplexes

We now turn to the dynamical and structural properties of mixed RNA-DNA helices which are characteristics of R-loops. The specific hybrids studied are r(CAG)4:d(CTG)4 and r(CUG)4:d(CAG)4. Two initial hybrid helices were constructed for each sequence, one with ideal A-RNA and the other with ideal B-DNA conformations, all with anti–anti glycosidic angles. Each hybrid was run for 1 μs, and for each sequence, there is convergence to a final helix that is independent of the initial conditions. Figure S18 plots the RMSD for these different runs and confirms their stability. RMSD values over the last 200 ns of these runs along with spreads are given in Table 3. The converged duplexes have characteristics intermediate to A-RNA and B-DNA. In particular, the RMSD values of r(CUG)4:d(CAG)4 are closer to A-RNA, while the values of r(CAG)4:d(CTG)4 are closer to B-DNA.

Table 3 Hybrid Duplexes Considered: Their Steps and Average RMSDa

 	steps	average RMSD	
r(CUG):d(CAG), initially A-RNA	CA/UG+AG/CU+GpC	1.77 (0.33)	
r(CUG):d(CAG), initially B-DNA	CA/UG+AG/CU+GpC	3.02 (0.46)	
r(CAG):d(CTG), initially A-RNA	CA/TG+AG/CT+GpC	1.96 (0.50)	
r(CAG):d(CTG), initially B-DNA	CA/TG+AG/CT+GpC	1.48 (0.39)	
a The base-pair steps formed by the hybrid duplexes, as well as average RMSD as obtained from the last 200 ns of a 1 μs simulation. For each case, the RMSD is measured with respect to the initial structure, e.g., for r(CUG):d(CTG) initially in its A-RNA form, the RMSD is measured with respect to the initial A-RNA structure, etc. Values in parentheses are the standard deviations of the RMSD distributions.

An analysis using PCA on the hybrid duplexes (SI S19) shows that they are more similar to their A-RNA form and the main dynamic motion for all of them is a combination of bending/unbending and winding/unwinding. The hybrid bases rA-dT and rU-dA form two WC type hydrogen bonds (rA_N1:dT_N3:H3, rA_N6:H61-dT_O4 and rU_O4-dA_N6:H61, rU_N3:H3-dA_N1) as shown in SI Figure S20. We did not observe Hoogsteen hydrogen bonds.

All of the hybrid duplexes have similar average stacking area and hydrogen bond number. The total energy from non-bonded interactions, which includes van der Waals and electrostatic interactions, is negative. Table 4 shows the total overlap area, number of hydrogen bonds, and non-bonded energy for each hybrid duplex. It is observed that the rA-dT hybrid duplex has a lower non-bonded energy than the rU-dA hybrid duplex, which is consistent with previous experimental results.103 This is likely due to stronger hydrogen bonds and better pi-stacking interactions in the rA-dT hybrid duplex compared to the rU-dA hybrid duplex, indicating that r(CAG):d(CTG) is more stable than r(CUG):d(CAG).

Table 4 Average Overlap Area, Number of Hydrogen Bonds (H-Bonds), and Non-Bonded Energies of the Hybrid Duplexesa

 	overlap area	number of H-bonds	van der Waals energy	electrostatic energy	
r(CUG):d(CAG)	24.78 (2.73)	22.99 (2.77)	–403.9 (8.01)	–270.7 (60.64)	
r(CAG):d(CTG)	23.44 (2.55)	23.27 (2.75)	–413.9 (8.11)	–505 (52.75)	
a All quoted values were obtained by averaging the quantities over the last 200 ns of the simulations. Only results for structures initially in the A-RNA form are quoted. The overlap area (in Å2) is defined as the overlap area between the polygons formed by the sugar rings on successive bases as computed via the 3DNA program. Only hydrogen bonds between complementary bases, e.g., 1-24, 2-23, were counted. The non-bonded van der Waals and electrostatic energies (both in kcal/mol) were obtained from the CPPTRAJ program via the “energy” command.

A summary of the analysis of the structural parameters of the hybrid duplexes is given in Figure 14a,b. As expected, the structural properties are intermediate between those of A-RNA and B-DNA, giving the structures greater flexibility than pure RNA or DNA duplexes. Roughly speaking, the twist and helical rise are similar for the hybrids, while inclination and bend angles show larger deviations and vary significantly between base-pair steps and appear to be negatively correlated with the helical rise. The properties of slide and Zp are two discriminatory parameters: slide smaller than −0.8 Å corresponds to A-form helices, while values larger than −0.8 Å are related to B-forms; for Zp, values larger than 1.5 Å corresponded to A-form helices, while values smaller than 0.5 Å is indicative of B-form helices.104 These two parameters show that the hybrids are closer to their A-form.

Figure 14 (a–c) Comparison of step parameters for hybrid and standard helices. Average twist, inclination, bend (), rise, slide, and Zp for middle base-pair steps of hybrid r(CUG)4:d(CAG)4, and r(CAG)4:d(CTG)4 helices and standard A-RNA and B-DNA. Colors are used as follows. Red: d(CTG:CAG)4 (B-DNA); orange: r(CUG:CAG)4 (A-RNA); blue: r(CUG)4:d(CAG)4 initially B-DNA; green: r(CUG)4:d(CAG)4 initially A-RNA; yellow: r(CAG)4:d(CTG)4 initially B-DNA; purple: r(CAG)4:d(CTG)4 initially A-RNA. Data were averaged over the last 200 ns. ‘X’ in the x-axis labels stands for either T, U, or A, according to the sequence.

We have also examined the ion distribution around the initially A-form hybrid structures. Some of the most prominent binding sites are shown in SI Figure S21a–d. Binding sites close to the six middle base-pairs involve dT_O4 (with time fraction ∼12%), rA_N7 (∼25%), dA_N7 (∼12%), and rU_O4 (∼24%) as well as weaker (less than 10%) binding sites including O2, OP1, and OP2 atoms. The bindings usually include a neighboring O6 atoms of G base belonging to an adjacent WC base pair. It is interesting to note that, unlike pure DNA/RNA, the predominant binding site for both dT and rU occurs in the major groove (O4). In the G-C hybrids bases, we observed the typical ionic interaction with major groove (O6) of G and close distance to C_N4, pretty similar to ionic interaction in B-DNA and A-RNA.93

Discussion

Although the mechanisms behind TREDs are complex and dependent on many factors, in the majority of cases, non-B-DNA conformations play a major role.40 Furthermore, mutant transcripts also contribute to the pathogenesis of TREDs through toxic RNA gain-of-function.6,16,17 Here, we report on an extensive MD investigation of CUG and CTG homoduplexes that are associated with DM1 diseases and their corresponding d(CTG):r(CAG) and d(CAG):r(CUG) hybrids that can form under bidirectional transcription; and GUC and GTC homoduplexes, not associated with any disease. We have also determined the preferred conformations of the U·U and T·T mismatches for the TR homoduplexes and investigated how these mismatches influence their overall conformation. Our main findings are as follows: 1. The global free energy minima associated with the U·U mismatches in the r(CUG) and r(GUC) homoduplexes correspond to anti–anti (-ac) conformations inside the helical core. Our results indicate that the U·U mismatches are very dynamic, exhibiting large fluctuations and with a number of hydrogen bonds that varies between zero and two. The next minima correspond to the symmetric anti–syn and syn–anti conformations which are about 6 kcal/mol above the global minimum. Finally, syn–syn conformations have an even higher free energy. Typical hydrogen bond conformations for the intrahelical mismatches are given in Table 1 and Figure 2a–c.

2. The global free energy minima associated with the T·T mismatches in the d(CTG) and d(GTC) homoduplexes correspond to anti–anti (-ap) conformations inside the helical core. The T·T mismatches are also quite dynamic, but the amplitudes of their fluctuations are smaller than for the corresponding U·U mismatches. As in the RNA case, the next minima corresponds to anti–syn and syn–anti conformations, followed by a higher syn–syn minimum. The most interesting feature is that free energy difference between the anti–syn and syn–anti minima and the absolute anti–anti minimum for d(GTC) is about half than that for d(CTG), and the free energy barriers for d(GTC) are also lower, a feature that is not observed for the RNA counterparts. There is also evidence for anti–anti minima with extruded mismatches and pseudo GpC steps in d(GTC). Table 1 and Figure 3a–c show typical hydrogen bond conformations for the mismatches.

3. We have carried out extensive MD simulations of all the mismatched duplexes starting from different mismatch conformations. Within these 1μs simulations, the mismatched duplexes remained stable in their initial anti–anti conformations. Duplexes with initial anti–syn or syn–syn conformations transitioned toward the anti–anti conformations although not all the mismatches reached their global minimum in the simulation time frame.

4. The U·U and T·T mismatches can flex between multiple conformations without significantly altering the global conformations of the helices. Fluctuations of helices around the first eigenvector direction in the PCA of the backbone show a coupling between bending and unwinding modes with more motion in the RNA helices. On a local level, opening, shear, and stretching are the dominant dynamical modes impacting the base-pair structure within the plane perpendicular to the helices. The RNA repeats are also stretched due to the widening of the major groove width and its undertwisting. The duplexes also exhibit a decrease in the inclination angle with respect to the canonical A-RNA or B-DNA structures. Local distortions cause overextension, unwinding, and bending.

5. We have characterized the neutralizing Na+ ion distribution surrounding the U·U and T·T mismatches. For U, the O4 atom is associated with the key binding site for all r(CUG) conformations, as well as for the r(GUC) anti–anti conformations. The other important ion binding site is associated with O2, which is observed in all conformations as having a high (relative) occupancy except for the anti–anti conformation. For d(CTG) and d(GTC), O2 represents the primary ion binding site, except for anti–syn and syn–syn in d(GTC) for which the O4 site dominates. For all duplexes, there is also ion binding associated with the backbone phosphate groups.

6. We have also extended our investigations to cover the important cases of the RNA–DNA hybrid duplexes that form from the expanding sequences, mainly r(CUG)4:d(CAG)4 and r(CAG)4:d(CTG)4 duplexes. We carried out 1 μs MD simulations of these helices, which are stable in anti–anti conformations. These hybrid helices are quite flexible and intermediate between pure A and B forms. Generally though, they are somewhat closer to the A-form, which may be attributed to the enhanced rigidity of the RNA strand over the DNA. These results are consistent with previous experimental and computational studies.105−108 Thermodynamic analysis shows that the most stable hybrids form between a purine-rich RNA transcript and the complementary pyrimidine-rich DNA template and that RNA duplexes are more stable than RNA(R-rich):DNA(Y-rich) hybrid duplexes, which in turn are more stable than DNA duplexes and DNA(R-rich):RNA(Y-rich) hybrid duplexes (which of the last two is more stable depends on the sequence).109 A detailed analysis of the helical and base-pair parameters indicate that the interplay between inclination and bend angles contributes to hybrid stretching and twist-stretch coupling (bending and unwinding movements). Slide and Zp are probably the best parameters to discriminate between the A and B-forms of the hybrid structures. According to our findings shown in Figure 14 as well as the RMSDs in SI Figure S18, the hybrid duplexes are somewhat closer to their A-RNA form.

7. Finally, with respect to the pathology of the related diseases, a very important question is which biophysical markers lead the CTG/CAG sequences—but not the GTC/GAC ones—to expand. This question has been addressed in a recent collaboration83 that combines single molecule FRET (smFRET) experiments and MD simulations to determine conformational stabilities and slipping dynamics for CAG, CTG, GAC, and GTC hairpins. Tetraloops are favored in CAG, CTG, and GTC, while GAC favors triloops. The evidence presented in that work suggests that the different loop stabilities have implications for intermediate structures that may form when DNA TR helices open. Thus, opposing hairpins in the (CAG)·(CTG) helix would have matched stability whereas opposing hairpins in a (GAC)·(GTC) duplex would have unmatched stability, introducing frustration in the non-expanding hairpins that in turn could anneal back to duplex DNA. The present work, combined with previous results,72 shows additional differences in the stems of the hairpins. First, there are different steps and energies associated with them, which have been reported in the literature, see for instance ref (110). In addition, our free energy maps show that in both GAC and GTC sequences, the symmetric anti–syn and syn–anti minima are deeper and wider than the corresponding CAG and CTG ones. Measured with respect to the global anti–anti minima, this second set of minima is approximately 0.9 and 2.0 kcal/mol for GAC and CAG, respectively; and they are approximately 3.5 and 7 kcal/mol for GTC and CTG, respectively. For homoduplexes with multiple mismatches, results shown in Figures S5 and S7 also support the notion that these secondary minima are longer-lived in GTC than in CTG. Thus, the anti–syn states are considerably more accessible in the non-expanding GNC sequences that in the expanding CNG sequences, which might further help the collapse of the non-expanding GNC hairpins when they interact with relevant proteins.

Conclusions

This work completes extensive studies carried out by our group and collaborators72,77,83 on CAG, GAC, CTG, and GTC repeats. In addition to exhaustive and careful evaluations of the hairpin structures and the homoduplexes, our combined work has found important structural and dynamical differences between the expanding and non-expanding sequences, thus identifying the biophysical factors associated with the expanding sequences that lead to disease. Among the many results in the present work, the fact that the mismatches in the GNC sequences can flip from anti–anti conformations to anti–syn conformations (which are long-lived) with much more ease than their CNG counterparts could thus provide another pathway in the repairing process, leading to the repair of anomalous structures and thus precluding the expansion in GNC sequences.

Supporting Information Available

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jpcb.3c03538.Additional simulation analysis and data (PDF)

Supplementary Material

jp3c03538_si_001.pdf

Author Contributions

§ A.F. and J.Q. contributed equally to this paper.

Author Contributions

Designed research: C.S. and C.R.; performed research: A.F., J.Q., and F.P., with greater part by A.F.; analyzed data: A.F., J.Q., and F.P., with greater part by A.F; and wrote paper: all authors.

The authors declare no competing financial interest.

Acknowledgments

The authors thank the NC State HPC Center for extensive computer support. This work has been supported by the National Institute of Health via Grant NIH-R01GM118508.
==== Refs
References

Oberle I. ; Rouseau F. ; Heitz D. ; Devys D. ; Zengerling S. ; Mandel J. Molecular-basis of the fragile-X syndrome and diagnostic applications. Am. J. Hum. Genet. 1991, 49 , 76.1829582
Giunti P. ; Sweeney M. G. ; Spadaro M. ; Jodice C. ; Novelletto A. ; Malaspina P. ; Frontali M. ; Harding A. E. The Trinucleotide Repeat Expansion on Chromosome 6p (SCA1) in Autosomal Dominant Cerebellar Ataxias. Brain 1994, 117 , 645–649. 10.1093/brain/117.4.645.7922453
Campuzano V. ; Montermini L. ; Molto M. ; Pianese L. ; Cossée M. ; Cavalcanti F. ; Monros E. ; Rodius F. ; Duclos F. ; Monticelli A. ; et al. Friedreich’s ataxia: Autosomal recessive disease caused by an intronic GAA triplet repeat expansion. Science 1996, 271 , 1423–1427. 10.1126/science.271.5254.1423.8596916
Ellegren H. Microsatellites: Simple sequences with complex evolution. Nat. Rev. Genet. 2004, 5 , 435–445. 10.1038/nrg1348.15153996
Mirkin S. M. DNA structures, repeat expansions and human hereditary disorders. Curr. Opin. Struct. Biol. 2006, 16 , 351–358. 10.1016/j.sbi.2006.05.004.16713248
Mirkin S. Expandable DNA repeats and human disease. Nature 2007, 447 , 932 10.1038/nature05977.17581576
Pearson C. ; Edamura K. ; Cleary J. Repeat instability: Mechanisms of dynamic mutations. Nat. Rev. Genet. 2005, 6 , 729–742. 10.1038/nrg1689.16205713
Wells R.D. ; Warren S. Genetic instabilities and neurological diseases; Academic Press: San Diego, CA, 1998.
Orr H. ; Zoghbi H. Trinucleotide repeat disorders. Annu. Rev. Neurosci. 2007, 30 , 575 10.1146/annurev.neuro.29.051605.113042.17417937
Pearson C. ; Sinden R. Slipped strand DNA (S-DNA and SI-DNA), trinucleotide repeat instability and mismatch repair: A short review. In Structure, Motion, Interaction and Expression of Biological Macromolecules, 10th Conversation in Biomolecular Stereodynamics Conference, SUNY Albany, JUN 17–21, 1997, 1998; Vol. 2 , pp 191–207.
Wells R. ; Dere R. ; Hebert M. ; Napierala M. ; Son L. Advances in mechanisms of genetic instability related to hereditary neurological diseases. Nucleic Acids Res. 2005, 33 , 3785–3798. 10.1093/nar/gki697.16006624
Kim J. C. ; Mirkin S. M. The balancing act of DNA repeat expansions. Curr. Opin. Genet. Dev. 2013, 23 , 280–288. 10.1016/j.gde.2013.04.009.23725800
Cleary J. ; Walsh D. ; Hofmeister J. ; Shankar G. ; Kuskowski M. ; Selkoe D. ; Ashe K. Natural oligomers of the amyloid-protein specifically disrupt cognitive function. Nat. Neurosci. 2005, 8 , 79–84. 10.1038/nn1372.15608634
Dion V. ; Wilson J. H. Instability and chromatin structure of expanded trinucleotide repeats. Trends Genet. 2009, 25 , 288–297. 10.1016/j.tig.2009.04.007.19540013
McMurray C. T. Hijacking of the mismatch repair system to cause CAG expansion and cell death in neurodegenerative disease. DNA Repair 2008, 7 , 1121–1134. 10.1016/j.dnarep.2008.03.013.18472310
Ranum L. P. W. ; Cooper T. A. RNA-mediated neuromuscular disorders. Annu. Rev. Neurosci. 2006, 29 , 259–277. 10.1146/annurev.neuro.29.051605.113014.16776586
Li L.-B. ; Bonini N. M. Roles of trinucleotide-repeat RNA in neurological disease and degeneration. Trends Neurosci. 2010, 33 , 292–298. 10.1016/j.tins.2010.03.004.20398949
Jin P. ; Zarnescu D. ; Zhang F. ; Pearson C. ; Lucchesi J. ; Moses K. ; Warren S. RNA-mediated neurodegeneration caused by the fragile X premutation rCGG repeats in Drosophila. Neuron 2003, 39 , 739–747. 10.1016/S0896-6273(03)00533-6.12948442
Jiang H. ; Mankodi A. ; Swanson M. ; Moxley R. ; Thornton C. Myotonic dystrophy type 1 is associated with nuclear foci of mutant RNA, sequestration of muscleblind proteins and deregulated alternative splicing in neurons. Hum. Mol. Genet. 2004, 13 , 3079–3088. 10.1093/hmg/ddh327.15496431
Daughters R. ; Tuttle D. ; Gao W. ; Ikeda Y. ; Moseley M. ; Ebner T. ; Swanson M. ; Ranum L. RNA Gain-of-Function in Spinocerebellar Ataxia Type 8. PLoS Genet. 2009, 5 , e1000600 10.1371/journal.pgen.1000600.19680539
Krzyzosiak W. ; Sobczak K. ; Wojciechowska M. ; Fiszer A. ; Mykowska A. ; Kozlowski P. Triplet repeat RNA structure and its role as pathogenic agent and therapeutic target. Nucleic Acids Res. 2012, 40 , 11–26. 10.1093/nar/gkr729.21908410
Campuzano V. ; Montermini L. ; Lutz Y. ; Cova L. ; Hindelang C. ; Jiralerspong S. ; Trottier Y. ; Kish S. ; Faucheux B. ; Trouillas P. ; et al. Frataxin is reduced in Friedreich ataxia patients and is associated with mitochondrial membranes. Hum. Mol. Genet. 1997, 6 , 1771–1780. 10.1093/hmg/6.11.1771.9302253
Kim E. ; Napierala M. ; Dent S. Hyperexpansion of GAA repeats affects post-initiation steps of FXN transcription in Friedreich’s ataxia. Nucleic Acids Res. 2011, 39 , 8366–8377. 10.1093/nar/gkr542.21745819
Kumari D. ; Biacsi R. ; Usdin K. Repeat expansion in intron 1 of the Frataxin gene reduces transcription initiation in Friedreich ataxia. FASEB J. 2011, 25 , 895.2 10.1096/fasebj.25.1_supplement.895.2.
Punga T. ; Buehler M. Long intronic GAA repeats causing Friedreich ataxia impede transcription elongation. EMBO Mol. Med. 2010, 2 , 120–129. 10.1002/emmm.201000064.20373285
Brook J. ; McCurrach J. ; Harley H. ; Buckler A. ; Church D. ; Aburatani H. ; Hunter K. ; Stanton V. ; Thirion J. ; Hudson T. ; Sohn R. ; Zemelman B. ; Snell R. G. ; Rundle S. A. ; Crow S. ; Davies J. ; Shelbourne P. ; Buxton J. ; Jones C. ; Juvonen V. ; Johnson K. ; Harper P. S. ; Shaw D. J. ; Housman D. E. Molecular basis of myotonic dystrophy: expansion of a trinucleotide (CTG) repeat at the 3′ end of a transcript encoding a protein kinase family member. Cell 1992, 68 , 799 10.1016/0092-8674(92)90154-5.1310900
Mahadevan M. ; Tsilfidis C. ; Sabourin L. ; Shutler G. ; Amemiya C. ; Jansen G. ; Neville C. ; Narang M. ; Barcelo J. ; O’Hoy K. ; et al. Myotonic dystrophy mutation: An unstable CTG repeat in the untranslated region of the gene. Science 1992, 255 , 1253 10.1126/science.1546325.1546325
Pettersson O. J. ; Aagaard L. ; Jensen T. G. ; Damgaard C. K. Molecular mechanisms in DM1 - a focus on foci. Nucleic Acids Res. 2015, 43 , 2433–2441. 10.1093/nar/gkv029.25605794
Shaw D. J. ; McCurrach M. ; Rundle S. A. ; Harley H. G. ; Crow S. R. ; Sohn R. ; Thirion J. P. ; Hamshere M. G. ; Buckler A. J. ; Harper P. S. ; Housman D. E. ; Brook J. D. Genomic organization and transcriptional units at the myotonic dystrophy locus. Genomics 1993, 18 , 673–679. 10.1016/S0888-7543(05)80372-6.7905855
Davis B. ; McCurrach M. ; Taneja K. ; Singer R. ; Housman D. Expansion of a CUG trinucleotide repeat in the 3′ untranslated region of myotonic dystrophy protein kinase transcripts results in nuclear retension of transcripts. Proc. Natl. Acad. Sci. U. S. A. 1997, 94 , 7388 10.1073/pnas.94.14.7388.9207101
Reddy K. ; Jenquin J. R. ; Cleary J. D. ; Berglund J. A. Mitigating RNA Toxicity in Myotonic Dystrophy using Small Molecules. Int. J. Mol. Sci. 2019, 20 , 4017 10.3390/ijms20164017.31426500
Philips A. ; Timchenko L. ; Cooper T. Disruption of splicing regulated by a CUG-binding protein in myotonic dystrophy. Science 1998, 280 , 737–741. 10.1126/science.280.5364.737.9563950
Timchenko A. ; Cai Z. ; Welm A. ; Reddy S. ; Ashizawa T. ; Timchenko L. RNA CUG repeats sequester CUGBP1 and alter protein levels and activity of CUGBP1. J. Biol. Chem. 2001, 276 , 7820 10.1074/jbc.M005960200.11124939
Savkur R. ; Philips A. ; Cooper T. Aberrant regulation of insulin receptor alternative splicing is associated with insulin resistance in myotonic dystrophy. Nat. Genet. 2001, 29 , 40–47. 10.1038/ng704.11528389
Ho T. ; Charlet B. ; Poulos M. ; Singh G. ; Swanson M. ; Cooper T. Muscleblind proteins regulate alternative splicing. EMBO J. 2004, 23 , 3103 10.1038/sj.emboj.7600300.15257297
Mykowska A. ; Sobczak K. ; Wojciechowska M. ; Kozlowski P. ; Krzyzosiak W. CAG repeats mimic CUG repeats in the misregulation of alternative splicing. Nucleic Acids Res. 2011, 39 , 8938 10.1093/nar/gkr608.21795378
Lin Y. ; Dent S. Y. ; Wilson J. H. ; Wells R. D. ; Napierala M. R loops stimulate genetic instability of CTG.CAG repeats. Proc. Natl. Acad. Sci. U. S. A. 2010, 107 , 692–697. 10.1073/pnas.0909740107.20080737
Reddy K. ; Tam M. ; Bowater R. P. ; Barber M. ; Tomlinson M. ; Nichol Edamura K. ; Wang Y. H. ; Pearson C. E. Determinants of R-loop formation at convergent bidirectionally transcribed trinucleotide repeats. Nucleic Acids Res. 2011, 39 , 1749–1762. 10.1093/nar/gkq935.21051337
Reddy K. ; Schmidt M. H. ; Geist J. M. ; Thakkar N. P. ; Panigrahi G. B. ; Wang Y. H. ; Pearson C. E. Processing of double-R-loops in (CAG)· (CTG) and C9orf72 (GGGGCC)· (GGCCCC) repeats causes instability. Nucleic Acids Res. 2014, 42 , 10473–10487. 10.1093/nar/gku658.25147206
McMurray C. DNA secondary structure: A common and causative factor for expansion in human disease. Proc. Natl. Acad. Sci. U. S. A. 1999, 96 , 1823–1825. 10.1073/pnas.96.5.1823.10051552
Michalowski S. ; Miller J. W. ; Urbinati C. R. ; Paliouras M. ; Swanson M. S. ; Griffith J. Visualization of double-stranded RNAs from the myotonic dystrophy protein kinase gene and interactions with CUG-binding protein. Nucleic Acids Res. 1999, 27 , 3534–3542. 10.1093/nar/27.17.3534.10446244
Sato Y. ; Kimura K. ; Takenaka A. Crystallization and preliminary X-ray analysis of RNA oligomers containing CUG repeats which induce type 1 myotonic dystrophy. Nucleic Acids Symp. Ser. 2004, 117–118. 10.1093/nass/48.1.117.
Mooers B. H. ; Logue J. S. ; Berglund J. A. The structural basis of myotonic dystrophy from the crystal structure of CUG repeats. Proc. Natl. Acad. Sci. U. S. A. 2005, 102 , 16626–16631. 10.1073/pnas.0505873102.16269545
Kumar A. ; Park H. ; Fang P. ; Parkesh R. ; Guo M. ; Nettles K. W. ; Disney M. D. Myotonic dystrophy type 1 RNA crystal structures reveal heterogeneous 1× 1 nucleotide UU internal loop conformations. Biochemistry 2011, 50 , 9928–9935. 10.1021/bi2013068.21988728
Parkesh R. ; Fountain M. ; Disney M. D. NMR spectroscopy and molecular dynamics simulation of r(CCGCUGCGG)2 reveal a dynamic UU internal loop found in myotonic dystrophy type 1. Biochemistry 2011, 50 , 599–601. 10.1021/bi101896j.21204525
Tamjar J. ; Katorcha E. ; Popov A. ; Malinina L. Structural dynamics of double-helical RNAs composed of CUG/CUG- and CUG/CGG-repeats. J. Biomol. Struct. Dyn. 2012, 30 , 505–523. 10.1080/07391102.2012.687517.22731704
Mariappan S. V. ; Garcoa A. E. ; Gupta G. Structure and dynamics of the DNA hairpins formed by tandemly repeated CTG triplets associated with myotonic dystrophy. Nucleic Acids Res. 1996, 24 , 775–783. 10.1093/nar/24.4.775.8604323
Gacy A. M. ; Goellner G. ; Juranic N. ; Macura S. ; McMurray C. T. Trinucleotide repeats that expand in human disease form hairpin structures in vitro. Cell 1995, 81 , 533–540. 10.1016/0092-8674(95)90074-8.7758107
Chi L. M. ; Lam S. L. Structural roles of CTG repeats in slippage expansion during DNA replication. Nucleic Acids Res. 2005, 33 , 1604–1617. 10.1093/nar/gki307.15767285
Petruska J. ; Arnheim N. ; Goodman M. Stability of intrastrand hairpin structures formed by the CAG/CTG class of DNA triplet repeats associated with neurological diseases. Nucleic Acids Res. 1996, 24 , 1992–1998. 10.1093/nar/24.11.1992.8668527
Chastain P. ; Sinden R. CTG repeats associated with human genetic disease are inherently flexible. J. Mol. Biol. 1998, 275 , 405–411. 10.1006/jmbi.1997.1502.9466918
Bacolla A. ; Gellibolian R. ; Shimizu M. ; Amirhaeri S. ; Kang S. ; Ohshima K. ; Larson J. ; Harvey S. ; Stollar B. ; Wells R. Flexible DNA: Genetically unstable CTG center dot CAG and CGG center dot CCG from human hereditary neuromuscular disease genes. J. Biol. Chem. 1997, 272 , 16783–16792. 10.1074/jbc.272.27.16783.9201983
Mitas M. ; Yu A. ; Dill J. ; Kamp T. ; Chambers E. ; Haworth I. Hairpin properties of single-stranded-DNA containing a GC-rich triplet repeat - (CTG)(15). Nucleic Acids Res. 1995, 23 , 1050–1059. 10.1093/nar/23.6.1050.7731793
Mariappan S. V. S. ; Garcia A. E. ; Gupta G. Structure and dynamics of the DNA hairpins formed by tandemly repeated CTG triplets associated with myotonic dystrophy. Nucleic Acids Res. 1996, 24 , 775–783. 10.1093/nar/24.4.775.8604323
Braida C. ; Stefanatos R. K. ; Adam B. ; Mahajan N. ; Smeets H. J. ; Niel F. ; Goizet C. ; Arveiler B. ; Koenig M. ; Lagier-Tourenne C. ; Mandel J. L. ; Faber C. G. ; de Die-Smulders C. E. M. ; Spaans F. ; Monckton D. G. Variant CCG and GGC Repeats within the CTG Expansion Dramatically Modify Mutational Dynamics and Likely Contribute Toward Unusual Symptoms in Some Myotonic Dystrophy Type 1 Patients. Hum. Mol. Genet. 2010, 19 , 1399–1412. 10.1093/hmg/ddq015.20080938
Timchenko L. ; Timchenko N. ; Caskey C. ; Roberts R. Novel proteins with binding specificity for DNA CTG repeats and RNA CUG repeats: Implications for myotonic dystrophy. Hum. Mol. Genet. 1996, 5 , 115–121. 10.1093/hmg/5.1.115.8789448
Hou M. ; Robinson H. ; Gao Y. ; Wang A. H. Crystal structure of actinomycin D bound to the CTG triplet repeat sequences linked to neurological diseases. Nucleic Acids Res. 2002, 30 , 4910–4917. 10.1093/nar/gkf619.12433994
Kiliszek A. ; Banaszak K. ; Dauter Z. ; Rypniewski W. The first crystal structures of RNA-PNA duplexes and a PNA-PNA duplex containing mismatches-toward anti-sense therapy against TREDs. Nucleic Acids Res. 2016, 44 , 1937–1943. 10.1093/nar/gkv1513.26717983
Coonrod L. A. ; Lohman J. R. ; Berglund J. A. Utilizing the GAAA tetraloop/receptor to facilitate crystal packing and determination of the structure of a CUG RNA helix. Biochemistry 2012, 51 , 8330–8337. 10.1021/bi300829w.23025897
Kiliszek A. ; Kierzek R. ; Krzyzosiak W. J. ; Rypniewski W. Structural insights into CUG repeats containing the ‘stretched U-U wobble’: implications for myotonic dystrophy. Nucleic Acids Res. 2009, 37 , 4149–4156. 10.1093/nar/gkp350.19433512
Yildirim I. ; Chakraborty D. ; Disney M. ; Wales D. ; Schatz G. Computational investigation of RNA CUG repeats responsible for myotonic dystrophy 1. J. Chem. Theory Comput. 2015, 11 , 4943 10.1021/acs.jctc.5b00728.26500461
Broda M. ; Kierzek E. ; Gdaniec Z. ; Kulinski T. ; Kierzek R. Thermodynamic stability of RNA structures formed by CNG trinucleotide repeats. Implication for prediction of RNA structure. Biochemistry 2005, 44 , 10873–10882. 10.1021/bi0502339.16086590
Yuan Y. ; Compton S. A. ; Sobczak K. ; Stenberg M. G. ; Thornton C. A. ; Griffith J. D. ; Swanson M. S. Muscleblind-like 1 interacts with RNA hairpins in splicing target and pathogenic RNAs. Nucleic Acids Res. 2007, 35 , 5474–5486. 10.1093/nar/gkm601.17702765
Kiliszek A. ; Kierzek R. ; Krzyzosiak W. J. ; Rypniewski W. Structural insights into CUG repeats containing the ‘stretched U-U wobble’: implications for myotonic dystrophy. Nucleic Acids Res. 2009, 37 , 4149–4156. 10.1093/nar/gkp350.19433512
Peyret N. ; Seneviratne P. A. ; Allawi H. T. ; SantaLucia J. Nearest-Neighbor Thermodynamics and NMR of DNA Sequences with Internal A· A, C· C, G· G, and T· T Mismatches. Biochemistry 1999, 38 , 3468–3477. 10.1021/bi9825091.10090733
Gervais V. ; Cognet J. A. ; Le Bret M. ; Sowers L. C. ; Fazakerley G. V. Solution structure of two mismatches A.A and T.T in the K-ras gene context by nuclear magnetic resonance and molecular dynamics. Eur. J. Biochem. 1995, 228 , 279–290. 10.1111/j.1432-1033.1995.00279.x.7705340
He G. ; Kwok C. K. ; Lam S. L. Preferential base pairing modes of T· T mismatches. FEBS Lett. 2011, 585 , 3953–3958. 10.1016/j.febslet.2011.10.044.22101148
Roy D. ; Yu K. ; Lieber M. R. Mechanism of R-loop formation at immunoglobulin class switch sequences. Mol. Cell. Biol. 2008, 28 , 50–60. 10.1128/MCB.01251-07.17954560
Roy D. ; Lieber M. R. G clustering is important for the initiation of transcription-induced R-loops in vitro, whereas high G density without clustering is sufficient thereafter. Mol. Cell. Biol. 2009, 29 , 3124–3133. 10.1128/MCB.00139-09.19307304
Sollier J. ; Cimprich K. A. Breaking bad: R-loops and genome integrity. Trends Cell Biol. 2015, 25 , 514–522. 10.1016/j.tcb.2015.05.003.26045257
Allison D. F. ; Wang G. G. R-loops: formation, function, and relevance to cell stress. Cell Stress 2019, 3 , 38–46. 10.15698/cst2019.02.175.31225499
Pan F. ; Man V. H. ; Roland C. ; Sagui C. Structure and Dynamics of DNA and RNA Double Helices of CAG and GAC Trinucleotide Repeats. Biophys. J. 2017, 113 , 19–36. 10.1016/j.bpj.2017.05.041.28700917
Zhang Y. ; Roland C. ; Sagui C. Structure and Dynamics of DNA and RNA Double Helices Obtained from the GGGGCC and CCCCGG Hexanucleotide Repeats that are the Hallmark of C9FTD/ALS Diseases. ACS Chem. Neurosci. 2017, 8 , 578–591. 10.1021/acschemneuro.6b00348.27933757
Zhang Y. ; Roland C. ; Sagui C. Structural and dynamical characterization of DNA and RNA quadruplexes obtained from the GGGGCC and GGGCCT hexanucleotide repeats associated with C9FTD/ALS and SCA36 diseases. ACS Chem. Neurosci. 2018, 9 , 1104 10.1021/acschemneuro.7b00476.29281254
Pan F. ; Zhang Y. ; Man V. H. ; Roland C. ; Sagui C. E-motif Formed by Extrahelical Cytosine Bases in DNA Homoduplexes of Trinucleotide and Hexanucleotide Repeats. Nucleic Acids Res. 2018, 46 , 942–955. 10.1093/nar/gkx1186.29190385
Pan F. ; Man V. ; Roland C. ; Sagui C. Structure and dynamics of DNA and RNA double helices obtained from the CCG and GGC trinucleotide repeats. J. Phys. Chem. B 2018, 122 , 4491 10.1021/acs.jpcb.8b01658.29617130
Xu P. ; Pan F. ; Roland C. ; Sagui C. ; Weninger K. Dynamics of strand slippage in DNA hairpins formed by CAG repeats: Roles of sequence parity and trinucleotide interrupts. Nucleic Acids Res. 2020, 48 , 2232–2245. 10.1093/nar/gkaa036.31974547
Zhang J. ; Fakharzadeh A. ; Pan F. ; Roland C. ; Sagui C. Atypical structures of GAA/TTC trinucleotide repeats underlying Friedreich’s ataxia: DNA triplexes and RNA/DNA hybrids. Nucleic Acids Res. 2020, 9899–9917. 10.1093/nar/gkaa665.32821947
Fakharzadeh A. ; Zhang J. ; Roland C. ; Sagui C. Novel eGZ-motif formed by regularly extruded guanine bases in a left-handed Z-DNA helix as a major motif behind CGG trinucleotide repeats. Nucleic Acids Res. 2022, 50 , 4860–4876. 10.1093/nar/gkac339.35536254
Zhang J. ; Fakharzadeh A. ; Pan F. ; Roland C. ; Sagui C. Construction of DNA/RNA Triplex Helices Based on GAA/TTC Trinucleotide Repeats. Bio-Protoc. 2021, 11 , e4155 10.21769/BioProtoc.4155.34692905
Pan F. ; Zhang Y. ; Xu P. ; Man V. H. ; Roland C. ; Weninger K. ; Sagui C. Molecular conformations and dynamics of nucleotide repeats associated with neurodegenerative diseases: double helices and CAG hairpin loops. Comput. Struct. Biotechnol. J. 2021, 19 , 2819–2832. 10.1016/j.csbj.2021.04.037.34093995
Zhang J. ; Fakharzadeh A. ; Roland C. ; Sagui C. RNA as a Major-Groove Ligand: RNA-RNA and RNA-DNA Triplexes Formed by GAA and UUC or TTC Sequences. ACS Omega 2022, 7 , 38728–38743. 10.1021/acsomega.2c04358.36340174
Xu P. ; Zhang J. ; Pan F. ; Mahn C. ; Roland C. ; Sagui C. ; Weninger K. Frustration Between Preferred States of Complementary Trinucleotide Repeat DNA Hairpins Anticorrelates with Expansion Disease Propensity. J. Mol. Biol. 2023, 435 , 168086 10.1016/j.jmb.2023.168086.37024008
Babin V. ; Roland C. ; Sagui C. Adaptively Biased Molecular Dynamics for Free Energy Calculations. J. Chem. Phys. 2008, 128 , 134101 10.1063/1.2844595.18397047
Case D. A. ; Ben-Shalom I. Y. ; Brozell S. R. ; Cerutti D. S. ; Cheatham T. E. III ; Cruzeiro V. W. D. ; Darden T. A. ; Duke R. E. ; Ghoreishi D. ; Gilson M. K. ; AMBER 2018; University of California: San Francisco CA, 2018.
Ivani I. ; Dans P. D. ; Noy A. ; Pérez A. ; Faustino I. ; Hopsital A. ; Walther J. ; Andrio P. ; Goñi R. ; Balaceanu A. ; et al. Parmbsc1: A Refined Force Field for DNA Simulations. Nat. Methods 2016, 13 , 55–58. 10.1038/nmeth.3658.26569599
Pérez A. ; March’an I. ; Svozil D. ; Sponer J. ; Cheatham T. E. III ; Laughton C. A. ; Orozco M. Refinement of the AMBER Force Field for Nucleic Acids: Improving the Description of α/γ Conformers. Biophys. J. 2007, 92 , 3817–3829. 10.1529/biophysj.106.097782.17351000
Zgarbová M. ; Šponer J. ; Otyepka M. ; Cheatham T. E. III ; Galindo-Murillo R. ; Jurečka P. Refinement of the Sugar-Phosphate Backbone Torsion Beta for AMBER Force Fields Improves the Description of Z- and B-DNA. J. Chem. Theory Comput. 2015, 11 , 5723–5736. 10.1021/acs.jctc.5b00716.26588601
Jorgensen W. L. ; Chandrasekhar J. ; Madura J. ; Klein M. L. Comparison of Simple Potential Functions for Simulating Liquid Water. J. Chem. Phys. 1983, 79 , 926–935. 10.1063/1.445869.
Joung I. S. ; Cheatham T. E. III Determination of Alkali and Halide Monovalent Ion Parameters for Use in Explicitly Solvated Biomolecular Simulations. J. Phys. Chem. B 2008, 112 , 9020–9041. 10.1021/jp8001614.18593145
Essmann U. ; Perera L. ; Berkowitz M. L. ; Darden T. ; Lee H. ; Pedersen L. G. A smooth particle mesh Ewald method. J. Chem. Phys. 1995, 103 , 8577–8593. 10.1063/1.470117.
Babin V. ; Karpusenka V. ; Moradi M. ; Roland C. ; Sagui C. Adaptively biased molecular dynamics: an umbrella sampling method with a time dependent potential. Int. J. Quantum Chem. 2009, 109 , 3666–3678. 10.1002/qua.22413.
Pan F. ; Roland C. ; Sagui C. Ion distribution around left- and right-handed DNA and RNA duplexes: a comparative study. Nucleic Acids Res. 2014, 42 , 13981 10.1093/nar/gku1107.25428372
Tiwary P. ; Parrinello M. A time-independent free energy estimator for metadynamics. J. Phys. Chem. B 2015, 119 , 736–742. 10.1021/jp504920s.25046020
Bussi G. ; Tribello G.A. Analyzing and biasing simulations with PLUMED. In Biomolecular Simulations: Methods and Protocols, 2019; Vol. 2022 , p 529.
Chen J. L. ; VanEtten D. ; Fountain M. ; Yildirim I. ; Disney M. Structure and dynamics of RNA repeat expansions that cause Huntington’s Disease and Myotonic Dystropy type 1. Biochemistry 2017, 56 , 3463 10.1021/acs.biochem.7b00252.28617590
Taghavi A. ; Yildirim I. Computational Investigation of Bending Properties of RNA AUUCU, CCUG, CAG, and CUG Repeat Expansions Associated With Neuromuscular Disorders. Front. Mol. Biosci. 2022, 9 , 830161 10.3389/fmolb.2022.830161.35480881
Lu X. J. ; Olson W. K. 3DNA: a software package for the analysis, rebuilding and visualization of three-dimensional nucleic acid structures. Nucleic Acids Res. 2003, 31 , 5108–5121. 10.1093/nar/gkg680.12930962
Satange R. ; Chang C. K. ; Hou M. H. A survey of recent unusual high-resolution DNA structures provoked by mismatches, repeats and ligand binding. Nucleic Acids Res. 2018, 46 , 6416–6434. 10.1093/nar/gky561.29945186
Gorin A. A. ; Zhurkin V. B. ; Olson W. K. B-DNA twisting correlates with base-pair morphology. J. Mol. Biol. 1995, 247 , 34–48. 10.1006/jmbi.1994.0120.7897660
Packer M. J. ; Dauncey M. P. ; Hunter C. A. Sequence-dependent DNA structure: tetranucleotide conformational maps. J. Mol. Biol. 2000, 295 , 85–103. 10.1006/jmbi.1999.3237.10623510
Lavery R. ; Moakher M. ; Maddocks J. H. ; Petkeviciute D. ; Zakrzewska K. Conformational analysis of nucleic acids revisited: Curves+. Nucleic Acids Res. 2009, 37 , 5917–5929. 10.1093/nar/gkp608.19625494
de Oliveira Martins E. ; Basílio Barbosa V. ; Weber G. DNA/RNA hybrid mesoscopic model shows strong stability dependence with deoxypyrimidine content and stacking interactions similar to RNA/RNA. Chem. Phys. Lett. 2019, 715 , 14–19. 10.1016/j.cplett.2018.11.015.
Lu X. J. ; Shakked Z. ; Olson W. K. A-form Conformational Motifs in Ligand-bound DNA Structures. J. Mol. Biol. 2000, 300 , 819–840. 10.1006/jmbi.2000.3690.10891271
Salazar M. ; Fedoroff O. Y. ; Miller J. M. ; Ribeiro N. S. ; Reid B. R. The DNA strand in DNA.RNA hybrid duplexes is neither B-form nor A-form in solution. Biochemistry 1993, 32 , 4207–4215. 10.1021/bi00067a007.7682844
Lesnik E. A. ; Freier S. M. Relative thermodynamic stability of DNA, RNA, and DNA:RNA hybrid duplexes: relationship with base composition and structure. Biochemistry 1995, 34 , 10807–10815. 10.1021/bi00034a013.7662660
Han G. W. ; Kopka M. L. ; Langs D. ; Sawaya M. R. ; Dickerson R. E. Crystal structure of an RNA.DNA hybrid reveals intermolecular intercalation: dimer formation by base-pair swapping. Proc. Natl. Acad. Sci. U. S. A. 2003, 100 , 9214–9219. 10.1073/pnas.1533326100.12872000
Liu J. H. ; Xi K. ; Zhang X. ; Bao L. ; Zhang X. ; Tan Z. J. Structural Flexibility of DNA-RNA Hybrid Duplex: Stretching and Twist-Stretch Coupling. Biophys. J. 2019, 117 , 74–86. 10.1016/j.bpj.2019.05.018.31164196
Sugimoto N. ; Nakano S. ; Katoh M. ; Matsumura A. ; Nakamuta H. ; Ohmichi T. ; Yoneyama M. ; Sasaki M. Thermodynamic parameters to predict stability of RNA/DNA hybrid duplexes. Biochemistry 1995, 34 , 11211–11216. 10.1021/bi00035a029.7545436
Hartenstine M. J. ; Goodman M. F. ; Petruska J. Base stacking and even/odd behavior of hairpin loops in DNA triplet repeat slippage and expansion with DNA polymerase. J. Biol. Chem. 2000, 275 , 18382–18390. 10.1074/jbc.275.24.18382.10849445
