
==== Front
J Med Chem
J Med Chem
jm
jmcmar
Journal of Medicinal Chemistry
0022-2623
1520-4804
American Chemical Society

37675804
10.1021/acs.jmedchem.3c01507
Viewpoint
Molecular Complexity: You Know It When You See It
https://orcid.org/0000-0002-6195-6976
Oprea Tudor I. *†‡
Bologa Cristian ‡
† Expert Systems Inc, 12760 High Bluff Dr #370, San Diego, California 92130, United States
‡ Department of Internal Medicine, University of New Mexico, MSC09-5025, Albuquerque, New Mexico 87131, United States
* Email: toprea@expertsystems.inc.
07 09 2023
28 09 2023
07 09 2024
66 18 1271012714
15 08 2023
© 2023 The Authors. Published by American Chemical Society
2023
The Authors
https://creativecommons.org/licenses/by-nc-nd/4.0/ Permits non-commercial access and re-use, provided that author attribution and integrity are maintained; but does not permit creation of adaptations or other derivative works (https://creativecommons.org/licenses/by-nc-nd/4.0/).
Molecular complexity (MC) lacks a universal definition, but various studies address it in contexts ranging from ligand–receptor interactions to DNA sequencing, with the overarching emphasis being its significance in synthetic organic chemistry and pharmaceutical research. Efforts to quantify MC in drug discovery have been numerous, but a unified approach remains challenging. Strategies based on graph theory, information theory, and substructural feature counts employed to gauge MC are often correlated to molecular weight (MW). Herbert Waldmann and his team introduced a new MC metric called the spacial score (SPS), which is based on factors like atom hybridization and stereoisomeric considerations. While SPS and its normalized version, nSPS, correlate with the natural product likeness score, they do not align with traditional chemical properties. We examined nSPS trends for approved drugs and found no significant changes in MC over eight decades, nor did nSPS capture drug innovation during that period. Furthermore, our analysis indicates that while the majority of approved drugs have an nSPS value between 10 and 20, this metric does not correlate with key drug properties like target bioactivity and oral bioavailability. Mirroring a chemist’s intuitive sense of chemical complexity, nSPS addresses the need for a precise empirical tool while a universal definition of MC remains elusive.

National Cancer Institute 10.13039/100000054 U24CA224370 document-id-old-9jm3c01507
document-id-new-14jm3c01507
ccc-price
==== Body
pmcThe term “molecular complexity” (MC) is not well-defined. “Molecular complexity” and “chemical complexity” are often used synonymously, but they can bear different interpretations. For example, Hann, Leach, and Harper discuss1 MC in molecular recognition by evaluating ligand patterns (e.g., shape, polarizability, and hydrophobicity) relevant to ligand–receptor interactions. Their approach aids in gauging the probability of ligand–receptor interactions. It is instrumental in contrasting lead compounds with drugs.2,3 Schuffenhauer et al. further refine4 the Hann model and demonstrate that, on average, biological activity tends to increase with higher MC by drawing on data from 160 assays and ∼19 200 compounds. On another front, Rujter, Scheffelaar, and Orru delve5 into MC from the standpoint of atom types, connectivity, and stereochemistry. Though not strictly mathematical, their perspective aligns more with the practical realms of diversity-oriented synthesis6 and biology-oriented synthesis.7 Such an approach also informs the total synthesis of medicinally important natural products,8 which hints that as synthetic organic chemistry evolves, “chemical complexity” might be simplified. Moving to a broader scale, Daley and Smith tackle9 MC in the context of DNA sequencing libraries by considering “the expected number of distinct molecules that can be observed in a given set of sequenced reads.”10 Perhaps one of the broader definitions of “chemical complexity” comes from Schmitt-Kopplin et al., who use complexity as a systems chemical analytics metric covering molecules, cells, ecosystems, and astronomy in living and abiotic systems.11 In this Viewpoint, our primary interest lies in the inherent chemical complexity pertinent to pharmaceutical research.

While many efforts have been made to quantify MC in the context of drug discovery—see Méndez-Lucio and Medina-Franco for a comprehensive survey12—a unified approach remains elusive. As Steven Bertz observed13 in 1981: “Synthetic chemists have been defining a ‘complex molecule’ in the way that many people define art: they know it when they see it. While the features which contribute to the complexity of a molecule have been discussed, no unified index has been formulated which takes into account the size, symmetry, branching, rings, multiple bonds, and heteroatoms characteristic of a complex molecule.” Bertz’s groundbreaking approach blended graph theory introduced by Bonchev and Trinajstic14 with principles of information theory, specifically Shannon entropy,15 to derive a “completely general” (cf. Bertz) MC index. However, Randić and Plavšić argue16 that graph theory and information theory deserve distinct recognition. They introduced MC terms that account for symmetry and other structural features. Using walk counts,17 particularly the total walk count (TWC), Rücker et al. assessed MC in the context of chemical synthesis.18 Whitlock19 integrated molecular size (on the basis of bond count) with “complexity” (counting rings, unsaturated bonds, and heteroatoms). This approach was refined by Barone and Chanon20 to provide better insights into ring structures and substitution patterns. We noted that TWC (in log10 format) and the Barone–Chanon MC metrics are highly correlated (R2 > 0.77) to molecular weight, MW.21 In response, we introduced SMCM, the synthetic and molecular complexity metric that considers hybridization, varied chiral centers (vicinal and geminal), spiro carbons, and more.21 By design, the SMCM is less correlated to MW (R2 = 0.535) than the previous metrics. Nevertheless, the expectation of a positive association between MC and MW is intuitive: given a set of diverse small molecules, an uptick in MW is likely to increase MC.

Writing in the Journal of Medicinal Chemistry, Herbert Waldmann and colleagues22 propose an innovative empirical metric for MC named the spacial score (SPS). This measure builds on established descriptors like the fraction of sp3 carbons, Fsp3,23 and the fraction of chiral centers.24 SPS combines an atom hybridization term, stereoisomeric factors, a nonaromatic ring term, and the number of heavy atom neighbors.22 Since SPS increases directly with MW, Waldmann et al. adjust SPS by the total heavy atom count in a given molecule, which results in the nSPS (normalized SPS). The authors offer a Github code (https://github.com/frog2000/Spacial-Score) for straightforward computation of both SPS and nSPS from molecular structures. Their in-depth analysis highlights a connection between SPS, nSPS, and Fsp3 and a correlation with the natural product likeness score.25 However, there’s no significant relationship between SPS parameters and more “conventional” chemical properties related to solubility, permeability, or molecular topology. Moreover, when juxtaposed with various MC metrics, including those from Bertz, Whitlock, and SMCM, SPS parameters show no correlation across various databases. Using the ChEMBL 30 database,26 increased potency is observed with increasing nSPS, which echoes Schuffenhauer et al.’s findings.4 Waldmann et al. also examined SPS and nSPS trends for small molecule drugs (SMDs) with a focus on those cataloged in DrugBank27 and approved by the FDA.28 Their analysis suggests that Fsp3 and nSPS values “do not appear to change appreciably over the years [1951–2021], and show no apparent increasing or decreasing trends.”22

Leveraging the publicly available nSPS source code, we independently evaluate nSPS trends for approved SMDs worldwide by drawing data from DrugCentral.29 On the basis of 4276 SMDs, the median nSPS is 15.654 (data not shown), with 66% of the drugs having nSPS ≤ 20 (Table 1). We then examine nSPS distribution for 1725 FDA-approved SMDs from an intellectual property (IP) perspective. Categorically, on the basis of IP rights, market exclusivity protections, and market accessibility,30 we distinguish OFP drugs, which are available in the market but have expired patents/exclusivities (N = 1038); ONP drugs, which are currently protected by patents/exclusivities and available in the market (N = 285); and OFM drugs, which are either discontinued or withdrawn (N = 402). In brief, ONP represents the newest drugs, OFP signifies somewhat older but still relevant therapeutics, while OFM captures outdated medicines. Observing no significant distribution shifts across these IP categories (Table 1), we conclude that nSPS does not capture novelty (as reflected by IP). To further drill into the temporal aspect, we simultaneously assessed the evolution of MW and MC in the drug discovery space, using nSPS as our metric. On the basis of the first approval dates (grouped by decade) for 1916 SMDs, it is evident that MW trends upward over time (Figure 1A). Meanwhile, nSPS maintains a distribution similar to that presented in Table 1, which indicates no marked growth in MC over the years (Figure 1B).

Table 1 Distribution of nSPS (Normalized Spacial Score) for Small Molecule Approved Drugsa

binned nSPS values	count	OFM	OFP	ONP	
<10	293	20	52	2	
10–15	1701	148	375	107	
15–20	857	67	187	86	
20–25	457	65	113	37	
25–30	267	28	66	19	
30–35	164	15	60	11	
35–40	123	8	50	5	
>40	414	51	135	18	
a OFM = off-market drugs; OFP = off-patent, on-market drugs; ONP = on-patent, on-market drugs.

Figure 1 Evolution of molecular weight (A) and the normalized spacial score (B) for 1916 approved drugs based on the first date of approval using data extracted from DrugCentral.

We also examine the relationship between MC and two critical SMD properties—target bioactivity and oral bioavailability (Figure 2). Using the nSPS metric, maximum target bioactivity values (negative log10) for each of the 1005 approved SMDs mirrors the distribution summarized in Table 1: most drugs have nSPS values between 10 and 20. This nSPS range encompasses 630 (63%) of the drugs with bioactivity data, of which 457 (45.5%) have maximum bioactivity of 100 nM or better, as detailed in Figure 2A (specific data omitted). Oral bioavailability, presented as a percentage of the administered dose (range from 0 to 100%), does not relate to nSPS either (Figure 2B). This finding also aligns with the distribution in Table 1. Narrowed to nSPS values between 10 and 20, we find 377 (61.7%) of the 611 drugs with bioavailability data. Among these, 292 (47.8%) have an oral bioavailability of 20% or higher (data not shown). On the basis of this analysis, it is plausible to suggest that such bioactivity and bioavailability enrichment might not be consistently observed in chemical databases that lack therapeutic compounds.

Figure 2 (A) Distribution of maximum bioactivity and nSPS for 1005 approved drugs; (B) distribution of oral bioavailability and nSPS for 611 approved drugs. Data from DrugCentral; nSPS, normalized spacial score. The maximum value was retained for drugs with bioactivity values across multiple targets.

Our independent analysis for approved SMDs from DrugCentral found no significant change in nSPS over time. Additionally, nSPS did not correlate significantly with IP novelty or with two important drug properties: target bioactivity and oral bioavailability. The nSPS metric may be the first MC measure that does not correlate with size, “traditional” cheminformatic descriptors, permeability, or bioactivity and encodes molecular complexity only. It further differentiates between the distribution of natural products versus synthetic compounds.22 Waldmann et al. write that nSPS “appears to reflect more the chemist’s intuitive assessment of complexity regarding spacial arrangements in molecules.” Medicinal chemists, pharmacologists, and the broader scientific community may struggle to find a universally accepted definition or absolute measure for MC. However, nSPS has emerged as a highly effective empirical tool by offering a precise measure that fills this gap.

This publication was supported in part by NIH grant U24CA224370.
==== Refs
References

Hann M. M. ; Leach A. R. ; Harper G. Molecular complexity and its impact on the probability of finding leads for drug discovery. J. Chem. Inf. Comput. Sci. 2001, 41 (3 ), 856–864. 10.1021/ci000403i.11410068
Teague S. J. ; Davis A. M. ; Leeson P. D. ; Oprea T. The Design of Leadlike Combinatorial Libraries. Angew. Chem., Int. Ed. 1999, 38 (24 ), 3743–3748. 10.1002/(SICI)1521-3773(19991216)38:24<3743::AID-ANIE3743>3.0.CO;2-U.
Oprea T. I. ; Davis A. M. ; Teague S. J. ; Leeson P. D. Is there a difference between leads and drugs? A historical perspective. J. Chem. Inf. Comput. Sci. 2001, 41 (5 ), 1308–1315. 10.1021/ci010366a.11604031
Schuffenhauer A. ; Brown N. ; Selzer P. ; Ertl P. ; Jacoby E. Relationships between molecular complexity, biological activity, and structural diversity. J. Chem. Inf. Model. 2006, 46 (2 ), 525–535. 10.1021/ci0503558.16562980
Ruijter E. ; Scheffelaar R. ; Orru R. V. A. Multicomponent reaction design in the quest for molecular complexity and diversity. Angew. Chem., Int. Ed. Engl. 2011, 50 (28 ), 6234–6246. 10.1002/anie.201006515.21710674
Schreiber S. L. Target-oriented and diversity-oriented organic synthesis in drug discovery. Science 2000, 287 (5460 ), 1964–1969. 10.1126/science.287.5460.1964.10720315
Nören-Müller A. ; Reis-Corrêa I. Jr ; Prinz H. ; Rosenbaum C. ; Saxena K. ; Schwalbe H. J. ; Vestweber D. ; Cagna G. ; Schunk S. ; Schwarz O. ; et al. Discovery of protein phosphatase inhibitor classes by biology-oriented synthesis. Proc. Natl. Acad. Sci. U. S. A. 2006, 103 (28 ), 10606–10611. 10.1073/pnas.0601490103.16809424
Nicolaou K. C. ; Hale C. R. H. ; Nilewski C. ; Ioannidou H. A. Constructing molecular complexity and diversity: total synthesis of natural products of biological and medicinal importance. Chem. Soc. Rev. 2012, 41 (15 ), 5185–5238. 10.1039/c2cs35116a.22743704
Daley T. ; Smith A. D. Predicting the molecular complexity of sequencing libraries. Nat. Methods 2013, 10 (4 ), 325–327. 10.1038/nmeth.2375.23435259
Chen Y. ; Negre N. ; Li Q. ; Mieczkowska J. O. ; Slattery M. ; Liu T. ; Zhang Y. ; Kim T.-K. ; He H. H. ; Zieba J. ; et al. Systematic evaluation of factors influencing ChIP-seq fidelity. Nat. Methods 2012, 9 (6 ), 609–614. 10.1038/nmeth.1985.22522655
Schmitt-Kopplin P. ; Hemmler D. ; Moritz F. ; Gougeon R. D. ; Lucio M. ; Meringer M. ; Müller C. ; Harir M. ; Hertkorn N. Systems chemical analytics: introduction to the challenges of chemical complexity analysis. Faraday Discuss. 2019, 218 , 9–28. 10.1039/C9FD00078J.31317165
Méndez-Lucio O. ; Medina-Franco J. L. The many roles of molecular complexity in drug discovery. Drug Discovery Today 2017, 22 (1 ), 120–126. 10.1016/j.drudis.2016.08.009.27575998
Bertz S. H. The first general index of molecular complexity. J. Am. Chem. Soc. 1981, 103 (12 ), 3599–3601. 10.1021/ja00402a071.
Bonchev D. ; Trinajstić N. Information theory, distance matrix, and molecular branching. J. Chem. Phys. 1977, 67 (10 ), 4517–4533. 10.1063/1.434593.
Shannon C. E. A mathematical theory of communication. Bell Syst. Technol. J. 1948, 27 (3 ), 379–423. 10.1002/j.1538-7305.1948.tb01338.x.
Randić M. ; Plavšić D. On the concept of molecular complexity. Croatica Chimica Acta 2002, 75 (1 ), 107–116.
Rücker G. ; Rücker C. Substructure, subgraph, and walk counts as measures of the complexity of graphs and molecules. J. Chem. Inf. Comput. Sci. 2001, 41 (6 ), 1457–1462. 10.1021/ci0100548.11749569
Rücker C. ; Rücker G. ; Bertz S. H. Organic synthesis – art or science. J. Chem. Inf. Comput. Sci. 2004, 44 (2 ), 378–386. 10.1021/ci030415e.15032515
Whitlock H. W. On the structure of total synthesis of complex natural products. J. Org. Chem. 1998, 63 (22 ), 7982–7989. 10.1021/jo9814546.
Barone R. ; Chanon M. A new and simple approach to chemical complexity. Application to the synthesis of natural products. J. Chem. Inf. Comput. Sci. 2001, 41 (2 ), 269–272. 10.1021/ci000145p.11277709
Allu T. K. ; Oprea T. I. Rapid evaluation of synthetic and molecular complexity for in silico chemistry. J. Chem. Inf. Model. 2005, 45 (5 ), 1237–1243. 10.1021/ci0501387.16180900
Krzyzanowski A. ; Pahl A. ; Grigalunas M. ; Waldmann H. Spacial score—A comprehensive topological indicator for small molecule complexity. J. Med. Chem. 2023, 10.1021/acs.jmedchem.3c00689.
Lovering F. ; Bikker J. ; Humblet C. Escape from flatland: Increasing saturation as an approach to improving clinical success. J. Med. Chem. 2009, 52 (21 ), 6752–6756. 10.1021/jm901241e.19827778
Clemons P. A. ; Bodycombe N. E. ; Carrinski H. A. ; Wilson J. A. ; Shamji A. F. ; Wagner B. K. ; Koehler A. N. ; Schreiber S. L. Small molecules of different origins have distinct distributions of structural complexity that correlate with protein-binding profiles. Proc. Natl. Acad. Sci. U. S. A. 2010, 107 (44 ), 18787–18792. 10.1073/pnas.1012741107.20956335
Grigalunas M. ; Burhop A. ; Zinken S. ; Pahl A. ; Gally J.-M. ; Wild N. ; Mantel Y. ; Sievers S. ; Foley D. J. ; Scheel R. ; et al. Natural product fragment combination to performance-diverse pseudo-natural products. Nat. Commun. 2021, 12 (1 ), 1883 10.1038/s41467-021-22174-4.33767198
Mendez D. ; Gaulton A. ; Bento A. P. ; Chambers J. ; De Veij M. ; Félix E. ; Magariños M. P. ; Mosquera J. F. ; Mutowo P. ; Nowotka M. ; et al. ChEMBL: towards direct deposition of bioassay data. Nucleic Acids Res. 2019, 47 (D1 ), D930–D940. 10.1093/nar/gky1075.30398643
Wishart D. S. ; Feunang Y. D. ; Guo A. C. ; Lo E. J. ; Marcu A. ; Grant J. R. ; Sajed T. ; Johnson D. ; Li C. ; Sayeeda Z. ; et al. DrugBank 5.0: a major update to the DrugBank database for 2018. Nucleic Acids Res. 2018, 46 (D1 ), D1074–D1082. 10.1093/nar/gkx1037.29126136
Scott K. A. ; Ropek N. ; Melillo B. ; Schreiber S. L. ; Cravatt B. F. ; Vinogradova E. V. Stereochemical diversity as a source of discovery in chemical biology. Curr. Res. Chem. Biol. 2022, 2 , 100028 10.1016/j.crchbi.2022.100028.
Avram S. ; Wilson T. B. ; Curpan R. ; Halip L. ; Borota A. ; Bora A. ; Bologa C. G. ; Holmes J. ; Knockel J. ; Yang J. J. ; et al. DrugCentral 2023 extends human clinical data and integrates veterinary drugs. Nucleic Acids Res. 2023, 51 (D1 ), D1276–D1287. 10.1093/nar/gkac1085.36484092
Avram S. ; Curpan R. ; Halip L. ; Bora A. ; Oprea T. I. Off-Patent Drug Repositioning. J. Chem. Inf. Model. 2020, 60 (12 ), 5746–5753. 10.1021/acs.jcim.0c00826.32877182
