
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

38565852
54655
10.1038/s41598-024-54655-z
Article
AI is a viable alternative to high throughput screening: a 318-target study
The Atomwise AIMS ProgramWallach Izhar 2
Bernard Denzil 2
Nguyen Kong 2
Ho Gregory 2
Morrison Adrian 2
Stecula Adrian 2
Rosnik Andreana 2
O’Sullivan Ann Marie 2
Davtyan Aram 2
Samudio Ben 2
Thomas Bill 2
Worley Brad 2
Butler Brittany 2
Laggner Christian 2
Thayer Desiree 2
Moharreri Ehsan 2
Friedland Greg 2
Truong Ha 2
van den Bedem Henry 2
Ng Ho Leung 2
Stafford Kate 2
Sarangapani Krishna 2
Giesler Kyle 2
Ngo Lien 2
Mysinger Michael 2
Ahmed Mostafa 2
Anthis Nicholas J. 2
Henriksen Niel 2
Gniewek Pawel 2
Eckert Sam 2
de Oliveira Saulo 2
Suterwala Shabbir 2
PrasadPrasad Srimukh Veccham Krishna 2
Shek Stefani 2
Contreras Stephanie 2
Hare Stephanie 2
Palazzo Teresa 2
O’Brien Terrence E. 2
Van Grack Tessa 2
Williams Tiffany 2
Chern Ting-Rong 2
Kenyon Victor 2
Lee Andreia H. 3
Cann Andrew B. 4
Bergman Bastiaan 5
Anderson Brandon M. 6
Cox Bryan D. 7
Warrington Jeffrey M. 8
Sorenson Jon M. 9
Goldenberg Joshua M. 10
Young Matthew A. 11
DeHaan Nicholas 12
Pemberton Ryan P. 13
Schroedl Stefan 14
Abramyan Tigran M. 1115
Gupta Tushita 16
Mysore Venkatesh 17
Presser Adam G. 18
Ferrando Adolfo A. 19
Andricopulo Adriano D. 20
Ghosh Agnidipta 21
Ayachi Aicha Gharbi 22
Mushtaq Aisha 23
Shaqra Ala M. 24
Toh Alan Kie Leong 25
Smrcka Alan V. 26
Ciccia Alberto 27
de Oliveira Aldo Sena 28
Sverzhinsky Aleksandr 29
de Sousa Alessandra Mara 30
Agoulnik Alexander I. 31
Kushnir Alexander 32
Freiberg Alexander N. 33
Statsyuk Alexander V. 34
Gingras Alexandre R. 35
Degterev Alexei 36
Tomilov Alexey 37
Vrielink Alice 38
Garaeva Alisa A. 39
Bryant-Friedrich Amanda 40
Caflisch Amedeo 41
Patel Amit K. 35
Rangarajan Amith Vikram 42
Matheeussen An 43
Battistoni Andrea 44
Caporali Andrea 45
Chini Andrea 46
Ilari Andrea 47
Mattevi Andrea 48
Foote Andrea Talbot 49
Trabocchi Andrea 50
Stahl Andreas 51
Herr Andrew B. 52
Berti Andrew 40
Freywald Andrew 53
Reidenbach Andrew G. 54
Lam Andrew 55
Cuddihy Andrew R. 56
White Andrew 57
Taglialatela Angelo 19
Ojha Anil K. 58
Cathcart Ann M. 59
Motyl Anna A. L. 45
Borowska Anna 39
D’Antuono Anna 60
Hirsch Anna K. H. 61
Porcelli Anna Maria 62
Minakova Anna 48
Montanaro Anna 60
Müller Anna 41
Fiorillo Annarita 63
Virtanen Anniina 64
O’Donoghue Anthony J. 35
Del Rio Flores Antonio 51
Garmendia Antonio E. 65
Pineda-Lucena Antonio 66
Panganiban Antonito T. 67
Samantha Ariela 38
Chatterjee Arnab K. 68
Haas Arthur L. 69
Paparella Ashleigh S. 21
John Ashley L. St. 70
Prince Ashutosh 71
ElSheikh Assmaa 72
Apfel Athena Marie 57
Colomba Audrey 73
O’Dea Austin 74
Diallo Bakary N’tji 75
Ribeiro Beatriz Murta Rezende Moraes 76
Bailey-Elkin Ben A. 77
Edelman Benjamin L. 78
Liou Benjamin 52
Perry Benjamin 79
Chua Benjamin Soon Kai 80
Kováts Benjámin 81
Englinger Bernhard 59
Balakrishnan Bijina 82
Gong Bin 33
Agianian Bogos 21
Pressly Brandon 37
Salas Brenda P. Medellin 83
Duggan Brendan M. 35
Geisbrecht Brian V. 84
Dymock Brian W. 85
Morten Brianna C. 85
Hammock Bruce D. 37
Mota Bruno Eduardo Fernandes 76
Dickinson Bryan C. 86
Fraser Cameron 87
Lempicki Camille 88
Novina Carl D. 89
Torner Carles 90
Ballatore Carlo 35
Bon Carlotta 91
Chapman Carly J. 92
Partch Carrie L. 93
Chaton Catherine T. 94
Huang Chang 65
Yang Chao-Yie 95
Kahler Charlene M. 38
Karan Charles 27
Keller Charles 96
Dieck Chelsea L. 97
Huimei Chen 70
Liu Chen 98
Peltier Cheryl 77
Mantri Chinmay Kumar 70
Kemet Chinyere Maat 55
Müller Christa E. 99
Weber Christian 100
Zeina Christina M. 59
Muli Christine S. 101
Morisseau Christophe 37
Alkan Cigdem 33
Reglero Clara 19
Loy Cody A. 101
Wilson Cornelia M. 102
Myhr Courtney 31
Arrigoni Cristina 48
Paulino Cristina 39
Santiago César 103
Luo Dahai 22
Tumes Damon J. 104
Keedy Daniel A. 105
Lawrence Daniel A. 57
Chen Daniel 106
Manor Danny 71
Trader Darci J. 101
Hildeman David A. 52
Drewry David H. 107
Dowling David J. 108
Hosfield David J. 86
Smith David M. 109
Moreira David 110
Siderovski David P. 111
Shum David 112
Krist David T. 113
Riches David W. H. 78
Ferraris Davide Maria 114
Anderson Deborah H. 115
Coombe Deirdre R. 116
Welsbie Derek S. 35
Hu Di 71
Ortiz Diana 117
Alramadhani Dina 118
Zhang Dingqiang 119
Chaudhuri Dipayan 82
Slotboom Dirk J. 39
Ronning Donald R. 120
Lee Donghan 121
Dirksen Dorian 122
Shoue Douglas A. 123
Zochodne Douglas William 124
Krishnamurthy Durga 52
Duncan Dustin 125
Glubb Dylan M. 92
Gelardi Edoardo Luigi Maria 126
Hsiao Edward C. 127
Lynn Edward G. 128
Silva Elany Barbosa 129
Aguilera Elena 130
Lenci Elena 50
Abraham Elena Theres 131
Lama Eleonora 62
Mameli Eleonora 45
Leung Elisa 125
Giles Ellie 102
Christensen Emily M. 132
Mason Emily R. 133
Petretto Enrico 70
Trakhtenberg Ephraim F. 134
Rubin Eric J. 18
Strauss Erick 135
Thompson Erik W. 25
Cione Erika 136
Lisabeth Erika Mathes 137
Fan Erkang 138
Kroon Erna Geessien 76
Jo Eunji 112
García-Cuesta Eva M. 103
Glukhov Evgenia 35
Gavathiotis Evripidis 21
Yu Fang 139
Xiang Fei 140
Leng Fenfei 141
Wang Feng 142
Ingoglia Filippo 82
van den Akker Focco 71
Borriello Francesco 143
Vizeacoumar Franco J. 144
Luh Frank 145
Buckner Frederick S. 138
Vizeacoumar Frederick S. 53
Bdira Fredj Ben 146
Svensson Fredrik 73
Rodriguez G. Marcela 147
Bognár Gabriella 81
Lembo Gaia 148
Zhang Gang 149
Dempsey Garrett 51
Eitzen Gary 150
Mayer Gaétan 151
Greene Geoffrey L. 86
Garcia George A. 57
Lukacs Gergely L. 152
Prikler Gergely 81
Parico Gian Carlo G. 93
Colotti Gianni 47
De Keulenaer Gilles 153
Cortopassi Gino 37
Roti Giovanni 60
Girolimetti Giulia 62
Fiermonte Giuseppe 154
Gasparre Giuseppe 155
Leuzzi Giuseppe 19
Dahal Gopal 156
Michlewski Gracjan 157158
Conn Graeme L. 159
Stuchbury Grant David 85
Bowman Gregory R. 160
Popowicz Grzegorz Maria 161
Veit Guido 152
de Souza Guilherme Eduardo 20
Akk Gustav 162
Caljon Guy 43
Alvarez Guzmán 163
Rucinski Gwennan 164
Lee Gyeongeun 112
Cildir Gökhan 165
Li Hai 27
Breton Hairol E. 166
Jafar-Nejad Hamed 167
Zhou Han 168
Moore Hannah P. 169
Tilford Hannah 164
Yuan Haynes 170
Shim Heesung 37
Wulff Heike 37
Hoppe Heinrich 75
Chaytow Helena 45
Tam Heng-Keat 171
Van Remmen Holly 172
Xu Hongyang 173
Debonsi Hosana Maria 174
Lieberman Howard B. 27
Jung Hoyoung 175
Fan Hua-Ying 176
Feng Hui 55
Zhou Hui 19
Kim Hyeong Jun 177
Greig Iain R. 178
Caliandro Ileana 179
Corvo Ileana 180
Arozarena Imanol 181
Mungrue Imran N. 182
Verhamme Ingrid M. 183
Qureshi Insaf Ahmed 184
Lotsaris Irina 185
Cakir Isin 57
Perry J. Jefferson P. 194
Kwiatkowski Jacek 85
Boorman Jacob 71
Ferreira Jacob 187
Fries Jacob 188
Kratz Jadel Müller 79
Miner Jaden 82
Siqueira-Neto Jair L. 35
Granneman James G. 189
Ng James 164
Shorter James 160
Voss Jan Hendrik 99
Gebauer Jan M. 131
Chuah Janelle 109
Mousa Jarrod J. 190
Maynes Jason T. 191
Evans Jay D. 192
Dickhout Jeffrey 193
MacKeigan Jeffrey P. 137
Jossart Jennifer N. 194
Zhou Jia 33
Lin Jiabei 160
Xu Jiake 195
Wang Jianghai 145
Zhu Jiaqi 196
Liao Jiayu 194
Xu Jingyi 194
Zhao Jinshi 197
Lin Jiusheng 198
Lee Jiyoun 199
Reis Joana 48
Stetefeld Joerg 77
Bruning John B. 200
Bruning John Burt 80
Coles John G. 201
Tanner John J. 166
Pascal John M. 29
So Jonathan 59
Pederick Jordan L. 80
Costoya Jose A. 110
Rayman Joseph B. 19
Maciag Joseph J. 52
Nasburg Joshua Alexander 37
Gruber Joshua J. 202
Finkelstein Joshua M. 55
Watkins Joshua 164
Rodríguez-Frade José Miguel 203
Arias Juan Antonio Sanchez 204
Lasarte Juan José 205
Oyarzabal Julen 204
Milosavljevic Julian 88
Cools Julie 153
Lescar Julien 22
Bogomolovas Julijus 35
Wang Jun 147
Kee Jung-Min 175
Kee Jung-Min 177
Liao Junzhuo 206
Sistla Jyothi C. 118
Abrahão Jônatas Santos 76
Sishtla Kamakshi 207
Francisco Karol R. 35
Hansen Kasper B. 208
Molyneaux Kathleen A. 71
Cunningham Kathryn A. 33
Martin Katie R. 137
Gadar Kavita 209
Ojo Kayode K. 138
Wong Keith S. 125
Wentworth Kelly L. 127
Lai Kent 82
Lobb Kevin A. 75
Hopkins Kevin M. 27
Parang Keykavous 210
Machaca Khaled 211
Pham Kien 98
Ghilarducci Kim 212
Sugamori Kim S. 125
McManus Kirk James 77
Musta Kirsikka 64
Faller Kiterie M. E. 45
Nagamori Kiyo 96
Mostert Konrad J. 135
Korotkov Konstantin V. 94
Liu Koting 213
Smith Kristiana S. 214
Sarosiek Kristopher 215
Rohde Kyle H. 216
Kim Kyu Kwang 217
Lee Kyung Hyeon 218
Pusztai Lajos 98
Lehtiö Lari 219
Haupt Larisa M. 25
Cowen Leah E. 125
Byrne Lee J. 102
Su Leila 145
Wert-Lamas Leon 89
Puchades-Carrasco Leonor 220
Chen Lifeng 86
Malkas Linda H. 186
Zhuo Ling 221
Hedstrom Lizbeth 222
Hedstrom Lizbeth 222
Walensky Loren D. 59
Antonelli Lorenzo 63
Iommarini Luisa 62
Whitesell Luke 125
Randall Lía M. 223
Fathallah M. Dahmani 224
Nagai Maira Harume 197
Kilkenny Mairi Louise 225
Ben-Johny Manu 19
Lussier Marc P. 212
Windisch Marc P. 112
Lolicato Marco 48
Lolli Marco Lucio 179
Vleminckx Margot 43
Caroleo Maria Cristina 226
Macias Maria J. 90
Valli Marilia 20
Barghash Marim M. 125
Mellado Mario 203
Tye Mark A. 227
Wilson Mark A. 198
Hannink Mark 228
Ashton Mark R. 85
Cerna Mark Vincent C.dela 121
Giorgis Marta 179
Safo Martin K. 118
Maurice Martin St. 229
McDowell Mary Ann 123
Pasquali Marzia 82
Mehedi Masfique 230
Serafim Mateus Sá Magalhães 76
Soellner Matthew B. 57
Alteen Matthew G. 231
Champion Matthew M. 123
Skorodinsky Maxim 232
O’Mara Megan L. 233
Bedi Mel 40
Rizzi Menico 114
Levin Michael 119
Mowat Michael 234
Jackson Michael R. 235
Paige Mikell 218
Al-Yozbaki Minnatallah 102
Giardini Miriam A. 129
Maksimainen Mirko M. 219
De Luise Monica 62
Hussain Muhammad Saddam 207
Christodoulides Myron 164
Stec Natalia 157
Zelinskaya Natalia 159
Van Pelt Natascha 43
Merrill Nathan M. 57
Singh Nathanael 105
Kootstra Neeltje A. 236
Singh Neeraj 237
Gandhi Neha S. 25
Chan Nei-Li 213
Trinh Nguyen Mai 22
Schneider Nicholas O. 229
Matovic Nick 85
Horstmann Nicola 238
Longo Nicola 82
Bharambe Nikhil 22
Rouzbeh Nirvan 208
Mahmoodi Niusha 21
Gumede Njabulo Joyfull 239
Anastasio Noelle C. 33
Khalaf Noureddine Ben 224
Rabal Obdulia 204
Kandror Olga 215
Escaffre Olivier 33
Silvennoinen Olli 64
Bishop Ozlem Tastan 75
Iglesias Pablo 110
Sobrado Pablo 240
Chuong Patrick 241
O’Connell Patrick 137
Martin-Malpartida Pau 90
Mellor Paul 53
Fish Paul V. 73
Moreira Paulo Otávio Lourenço 30
Zhou Pei 197
Liu Pengda 107
Liu Pengda 107
Wu Pengpeng 242
Agogo-Mawuli Percy 111
Jones Peter L. 243
Ngoi Peter 93
Toogood Peter 57
Ip Philbert 125
von Hundelshausen Philipp 100
Lee Pil H. 57
Rowswell-Turner Rachael B. 217
Balaña-Fouce Rafael 244
Rocha Rafael Eduardo Oliveira 76
Guido Rafael V. C. 20
Ferreira Rafaela Salgado 76
Agrawal Rajendra K. 58
Harijan Rajesh K. 21
Ramachandran Rajesh 245
Verma Rajkumar 246
Singh Rakesh K. 247
Tiwari Rakesh Kumar 248
Mazitschek Ralph 227
Koppisetti Rama K. 166
Dame Remus T. 146
Douville Renée N. 249
Austin Richard C. 193
Taylor Richard E. 123
Moore Richard G. 217
Ebright Richard H. 147
Angell Richard M. 73
Yan Riqiang 237
Kejriwal Rishabh 65
Batey Robert A. 125
Blelloch Robert 127
Vandenberg Robert J. 185
Hickey Robert J. 186
Kelm Robert J. Jr. 49
Lake Robert J. 176
Bradley Robert K. 250
Blumenthal Robert M. 106
Solano Roberto 46
Gierse Robin Matthias 251
Viola Ronald E. 156
McCarthy Ronan R. 209
Reguera Rosa Maria 244
Uribe Ruben Vazquez 252
do Monte-Neto Rubens Lima 30
Gorgoglione Ruggiero 154
Cullinane Ryan T. 222
Katyal Sachin 170
Hossain Sakib 105
Phadke Sameer 57
Shelburne Samuel A. 238
Geden Sandra E. 216
Johannsen Sandra 61
Wazir Sarah 219
Legare Scott 77
Landfear Scott M. 117
Radhakrishnan Senthil K. 118
Ammendola Serena 44
Dzhumaev Sergei 253
Seo Seung-Yong 140
Li Shan 142
Zhou Shan 167
Chu Shaoyou 133
Chauhan Shefali 254
Maruta Shinsaku 255256
Ashkar Shireen R. 57
Shyng Show-Ling 117
Conticello Silvestro G. 148256
Buroni Silvia 48
Garavaglia Silvia 114
White Simon J. 65
Zhu Siran 157158
Tsimbalyuk Sofiya 257
Chadni Somaia Haque 141
Byun Soo Young 112
Park Soonju 112
Xu Sophia Q. 258
Banerjee Sourav 259
Zahler Stefan 221
Espinoza Stefano 91
Gustincich Stefano 91
Sainas Stefano 179
Celano Stephanie L. 137
Capuzzi Stephen J. 107
Waggoner Stephen N. 52
Poirier Steve 260
Olson Steven H. 235
Marx Steven O. 261
Van Doren Steven R. 166
Sarilla Suryakala 183
Brady-Kalnay Susann M. 71
Dallman Sydney 230
Azeem Syeda Maryam 105
Teramoto Tadahisa 262
Mehlman Tamar 105
Swart Tarryn 75
Abaffy Tatjana 263
Akopian Tatos 215
Haikarainen Teemu 64
Moreda Teresa Lozano 264
Ikegami Tetsuro 33
Teixeira Thaiz Rodrigues 174
Jayasinghe Thilina D. 120
Gillingwater Thomas H. 45
Kampourakis Thomas 265
Richardson Timothy I. 207
Herdendorf Timothy J. 84
Kotzé Timothy J. 135
O’Meara Timothy R. 266
Corson Timothy W. 207
Hermle Tobias 88
Ogunwa Tomisin Happy 255
Lan Tong 86
Su Tong 228
Banjo Toshihiro 267
O’Mara Tracy A. 92
Chou Tristan 42
Chou Tsui-Fen 142
Baumann Ulrich 131
Desai Umesh R. 118
Pai Vaibhav P. 119
Thai Van Chi 38
Tandon Vasudha 259
Banerji Versha 77
Robinson Victoria L. 65
Gunasekharan Vignesh 168
Namasivayam Vigneshwaran 99
Segers Vincent F. M. 43
Maranda Vincent 53
Dolce Vincenza 136
Maltarollo Vinícius Gonçalves 76
Scoffone Viola Camilla 48
Woods Virgil A. 105
Ronchi Virginia Paola 268
Van Hung Le Vuong 269
Clayton W. Brent 101
Lowther W. Todd 270
Houry Walid A. 125
Li Wei 271
Tang Weiping 206
Zhang Wenjun 51
Van Voorhis Wesley C. 138
Donaldson William A. 229
Hahn William C. 59
Kerr William G. 272
Gerwick William H. 129
Bradshaw William J. 273
Foong Wuen Ee 274
Blanchet Xavier 275
Wu Xiaoyang 86
Lu Xin 123
Qi Xin 245
Xu Xin 84
Yu Xinfang 167
Qin Xingping 276
Wang Xingyou 222
Yuan Xinrui 95
Zhang Xu 277
Zhang Yan Jessie 83
Hu Yanmei 147
Aldhamen Yasser Ali 137
Chen Yicheng 71
Li Yihe 71
Sun Ying 52
Zhu Yini 123
Gupta Yogesh K. 278
Pérez-Pertejo Yolanda 244
Li Yong 167
Tang Young 65
He Yuan 40
Tse-Dinh Yuk-Ching 141
Sidorova Yulia A. 279
Yen Yun 145
Li Yunlong 280
Frangos Zachary J. 281
Chung Zara 22
Su Zhengchen 33
Wang Zhenghe 71
Zhang Zhiguo 27
Liu Zhongle 125
Inde Zintis 215
Artía Zoraima 163
Heifets Abraham 2
izhar@atomwise.com

1
1 San Fransico, CA USA
2 Atomwise Inc., San Fransico, USA
3 grid.417886.4 0000 0001 0657 5612 Amgen, Thousand Oaks, USA
4 https://ror.org/05wx9n238 grid.511328.c OpenAI, San Francisco, USA
5 Model Medicines, La Jolla, USA
6 Atomic.AI, San Francisco, USA
7 Edifice Health, Inc., San Mateo, USA
8 METiS Therapeutics, Cambridge, USA
9 https://ror.org/04gndp242 0000 0004 5899 3818 Genentech, San Mateo, USA
10 US Navy Medical Service Corps Officer (2300/1810D), San Mateo, USA
11 Totus Medicines, Inc., Emeryville, USA
12 https://ror.org/03tx9ss94 grid.421748.c 0000 0004 0460 2009 Cytokinetics, Inc., South San Francisco, USA
13 Nurix Therapeutics, San Francisco, USA
14 Amazon Alexa, Suite, USA
15 https://ror.org/0130frc33 grid.10698.36 0000 0001 2248 3208 The University of North Carolina at Chapel Hill Eshelman School of Pharmacy, Chapel Hill, USA
16 Refibered Inc., Cupertino, USA
17 https://ror.org/03jdj4y14 grid.451133.1 0000 0004 0458 4453 NVIDIA, Santa Clara, USA
18 grid.38142.3c 000000041936754X Harvard TH Chan School of Public Health, Boston, USA
19 https://ror.org/00hj8s172 grid.21729.3f 0000 0004 1936 8729 Columbia University, New York, USA
20 https://ror.org/036rp1748 grid.11899.38 0000 0004 1937 0722 University of São Paulo, São Paulo, Brazil
21 https://ror.org/05cf8a891 grid.251993.5 0000 0001 2179 1997 Albert Einstein College of Medicine, Bronx, USA
22 https://ror.org/02e7b5302 grid.59025.3b 0000 0001 2224 0361 Nanyang Technological University, Singapore, Singapore
23 https://ror.org/00cvxb145 grid.34477.33 0000 0001 2298 6657 University of Washington, Seattle, USA
24 https://ror.org/04ydmy275 grid.266685.9 0000 0004 0386 3207 Chan Medical School, University of Massachusetts, Worcester, USA
25 https://ror.org/03pnv4752 grid.1024.7 0000 0000 8915 0953 Queensland University of Technology, Brisbane, Australia
26 grid.214458.e 0000000086837370 University of Michigan Medical School, Ann Arbor, USA
27 https://ror.org/01esghr10 grid.239585.0 0000 0001 2285 2675 Columbia University Irving Medical Center, New York, USA
28 https://ror.org/041akq887 grid.411237.2 0000 0001 2188 7235 Universidade Federal de Santa Catarina, Florianópolis, Brazil
29 https://ror.org/0161xgx34 grid.14848.31 0000 0001 2104 2136 Université de Montréal, Montreal, Canada
30 grid.418068.3 0000 0001 0723 0931 Instituto René Rachou-Fundação Oswaldo Cruz/Fiocruz Minas, Belo Horizonte, Brazil
31 https://ror.org/02gz6gg07 grid.65456.34 0000 0001 2110 1845 Herbert Wertheim College of Medicine, Biomolecular Science Institute, Florida International University, Miami, USA
32 https://ror.org/005dvqh91 grid.240324.3 0000 0001 2109 4251 NYU Langone Health, New York, USA
33 https://ror.org/016tfm930 grid.176731.5 0000 0001 1547 9964 The University of Texas Medical Branch at Galveston, Galveston, USA
34 https://ror.org/048sx0r50 grid.266436.3 0000 0004 1569 9707 University of Houston, Galveston, USA
35 grid.266100.3 0000 0001 2107 4242 University of California, San Diego, USA
36 grid.429997.8 0000 0004 1936 7531 School of Medicine, Tufts University, Medford, USA
37 https://ror.org/05t99sp05 grid.468726.9 0000 0004 0486 2046 University of California, Davis, Davis, USA
38 https://ror.org/047272k79 grid.1012.2 0000 0004 1936 7910 University of Western Australia, Crawley, Australia
39 https://ror.org/012p63287 grid.4830.f 0000 0004 0407 1981 University of Groningen, Groningen, The Netherlands
40 https://ror.org/01070mq45 grid.254444.7 0000 0001 1456 7807 Wayne State University, Detroit, USA
41 https://ror.org/02crff812 grid.7400.3 0000 0004 1937 0650 University of Zurich, Zürich, Switzerland
42 https://ror.org/00f54p054 grid.168010.e 0000 0004 1936 8956 Stanford University, Stanford, USA
43 https://ror.org/008x57b05 grid.5284.b 0000 0001 0790 3681 University of Antwerp, Antwerp, Belgium
44 https://ror.org/02p77k626 grid.6530.0 0000 0001 2300 0941 University of Rome Tor Vergata, Rome, Italy
45 https://ror.org/01nrxwf90 grid.4305.2 0000 0004 1936 7988 University of Edinburgh, Edinburgh, UK
46 grid.4711.3 0000 0001 2183 4846 Department of Plant Molecular Genetics, Centro Nacional de Biotecnología, Consejo Superior de Investigaciones Científicas (CNB-CSIC), Madrid, Spain
47 grid.5326.2 0000 0001 1940 4177 CNR (Italian National Research Council), Rome, Italy
48 https://ror.org/00s6t1f81 grid.8982.b 0000 0004 1762 5736 University of Pavia, Pavia, Italy
49 https://ror.org/0155zta11 grid.59062.38 0000 0004 1936 7689 University of Vermont, Burlington, USA
50 https://ror.org/04jr1s763 grid.8404.8 0000 0004 1757 2304 University of Florence, Florence, Italy
51 https://ror.org/05t99sp05 grid.468726.9 0000 0004 0486 2046 University of California, Berkeley, Berkeley, USA
52 https://ror.org/01hcyya48 grid.239573.9 0000 0000 9025 8099 Cincinnati Children’s Hospital Medical Center, Cincinnati, USA
53 https://ror.org/010x8gc63 grid.25152.31 0000 0001 2154 235X University of Saskatchewan, Saskatoon, Canada
54 https://ror.org/05a0ya142 grid.66859.34 0000 0004 0546 1623 Broad Institute of MIT and Harvard, Cambridge, USA
55 https://ror.org/05qwgg493 grid.189504.1 0000 0004 1936 7558 Boston University, Boston, USA
56 grid.419404.c 0000 0001 0701 0170 CancerCare Manitoba Research Institute, Winnipeg, Canada
57 https://ror.org/00jmfr291 grid.214458.e 0000 0004 1936 7347 University of Michigan, Ann Arbor, USA
58 grid.465543.5 0000 0004 0435 9002 Wadsworth Center, New York State Department of Health and University at Albany, Albany, USA
59 https://ror.org/02jzgtq86 grid.65499.37 0000 0001 2106 9910 Dana-Farber Cancer Institute, Boston, USA
60 https://ror.org/02k7wn190 grid.10383.39 0000 0004 1758 0937 University of Parma, Parma, Italy
61 https://ror.org/042dsac10 grid.461899.b Helmholtz Institute for Pharmaceutical Research Saarland, Saarbrücken, Germany
62 https://ror.org/01111rn36 grid.6292.f 0000 0004 1757 1758 University of Bologna, Bologna, Italy
63 grid.7841.a Sapienza University of Rome, Rome, Italy
64 https://ror.org/033003e23 grid.502801.e 0000 0001 2314 6254 Tampere University, Tampere, Finland
65 https://ror.org/02der9h97 grid.63054.34 0000 0001 0860 4915 University of Connecticut, Storrs, USA
66 https://ror.org/02rxc7m23 grid.5924.a 0000 0004 1937 0271 Centro de Investigación Médica Aplicada, Universidad de Navarra, Pamplona, Spain
67 https://ror.org/04vmvtb21 grid.265219.b 0000 0001 2217 8588 Tulane National Primate Research Center, Tulane University, Covington, USA
68 grid.214007.0 0000000122199231 Scripps Research, San Diego, USA
69 https://ror.org/05ect4e57 grid.64337.35 0000 0001 0662 7451 Louisiana State University School of Medicine, New Orleans, USA
70 https://ror.org/02j1m6098 grid.428397.3 0000 0004 0385 0924 Duke-NUS Medical School, Singapore, Singapore
71 https://ror.org/051fd9666 grid.67105.35 0000 0001 2164 3847 Case Western Reserve University, Cleveland, USA
72 grid.412258.8 0000 0000 9477 7793 Oregon Health and Science University and Tanta University in Tanta, Tanta, Egypt
73 https://ror.org/02jx3x895 grid.83440.3b 0000 0001 2190 1201 University College London, London, UK
74 https://ror.org/01p7jjy08 grid.262962.b 0000 0004 1936 9342 Saint Louis University, St. Louis, USA
75 https://ror.org/016sewp10 grid.91354.3a 0000 0001 2364 1300 Rhodes University, Makhanda, South Africa
76 https://ror.org/0176yjw32 grid.8430.f 0000 0001 2181 4888 Universidade Federal de Minas Gerais (UFMG), Belo Horizonte, Brazil
77 https://ror.org/02gfys938 grid.21613.37 0000 0004 1936 9609 University of Manitoba, Winnipeg, Canada
78 https://ror.org/016z2bp30 grid.240341.0 0000 0004 0396 0728 National Jewish Health, Denver, USA
79 https://ror.org/022mz6y25 grid.428391.5 0000 0004 0618 1092 Drugs for Neglected Diseases Initiative (DNDi), Geneva, Switzerland
80 https://ror.org/00892tw58 grid.1010.0 0000 0004 1936 7304 The University of Adelaide, Adelaide, Australia
81 Mcule, Budapest, Hungary
82 https://ror.org/03r0ha626 grid.223827.e 0000 0001 2193 0096 University of Utah, Salt Lake City, USA
83 https://ror.org/00hj54h04 grid.89336.37 0000 0004 1936 9924 The University of Texas at Austin, Austin, USA
84 https://ror.org/05p1j8758 grid.36567.31 0000 0001 0737 1259 Kansas State University, Manhattan, USA
85 UniQuest Pty Ltd, St Lucia, Australia
86 https://ror.org/024mw5h28 grid.170205.1 0000 0004 1936 7822 University of Chicago, Chicago, USA
87 https://ror.org/03vek6s52 grid.38142.3c 0000 0004 1936 754X Harvard University, Cambridge, USA
88 https://ror.org/0245cg223 grid.5963.9 0000 0004 0491 7203 University of Freiburg, Freiburg Im Breisgau, Germany
89 grid.65499.37 0000 0001 2106 9910 Dana-Farber Cancer Institute and Harvard Medical School, Boston, USA
90 https://ror.org/01z1gye03 grid.7722.0 0000 0001 1811 6966 IRB Barcelona, Barcelona, Spain
91 https://ror.org/042t93s57 grid.25786.3e 0000 0004 1764 2907 Istituto Italiano Di Tecnologia, Genoa, Italy
92 https://ror.org/004y8wk30 grid.1049.c 0000 0001 2294 1395 QIMR Berghofer Medical Research Institute, Herston, Australia
93 https://ror.org/05t99sp05 grid.468726.9 0000 0004 0486 2046 University of California, Santa Cruz, Santa Cruz, USA
94 https://ror.org/02k3smh20 grid.266539.d 0000 0004 1936 8438 University of Kentucky, Lexington, USA
95 https://ror.org/0011qv509 grid.267301.1 0000 0004 0386 9246 University of Tennessee Health Science Center, Memphis, USA
96 https://ror.org/04netx779 grid.468147.8 Children’s Cancer Therapy Development Institute, Beaverton, USA
97 https://ror.org/01esghr10 grid.239585.0 0000 0001 2285 2675 Columbia University Medical Center, New York, USA
98 grid.47100.32 0000000419368710 Yale School of Medicine, New Haven, USA
99 https://ror.org/041nas322 grid.10388.32 0000 0001 2240 3300 University of Bonn, Bonn, Germany
100 https://ror.org/05591te55 grid.5252.0 0000 0004 1936 973X Ludwig-Maximilians-Universität München, Munich, Germany
101 https://ror.org/02dqehb95 grid.169077.e 0000 0004 1937 2197 Purdue University, West Lafayette, USA
102 https://ror.org/0489ggv38 grid.127050.1 0000 0001 0249 951X Canterbury Christ Church University, Canterbury, UK
103 grid.428469.5 0000 0004 1794 1018 National Centre for Biotechnology (CNB-CSIC), Madrid, Spain
104 grid.1026.5 0000 0000 8994 5086 University of South Australia and SA Pathology, Adelaide, Australia
105 https://ror.org/01gdjt538 grid.456297.b 0000 0004 5895 2063 CUNY Advanced Science Research Center, New York, USA
106 https://ror.org/01pbdzh19 grid.267337.4 0000 0001 2184 944X The University of Toledo, Toledo, USA
107 https://ror.org/0130frc33 grid.10698.36 0000 0001 2248 3208 University of North Carolina at Chapel Hill, Chapel Hill, USA
108 https://ror.org/00dvg7y05 grid.2515.3 0000 0004 0378 8438 Boston Children’s Hospital and Harvard Medical School, Boston, USA
109 https://ror.org/011vxgd24 grid.268154.c 0000 0001 2156 6140 West Virginia University, Morgantown, USA
110 https://ror.org/030eybx10 grid.11794.3a 0000 0001 0941 0645 Universidade de Santiago de Compostela, Santiago, Spain
111 https://ror.org/05msxaq47 grid.266871.c 0000 0000 9765 6057 University of North Texas Health Science Center at Fort Worth, Fort Worth, USA
112 https://ror.org/04t0zhb48 grid.418549.5 0000 0004 0494 4850 Institut Pasteur Korea, Seongnam, South Korea
113 grid.185648.6 0000 0001 2175 0319 Carle Illinois College of Medicine, Urbana, USA
114 grid.16563.37 0000000121663741 Università del Piemonte Orientale, Vercelli, Italy
115 https://ror.org/00e1nmf62 grid.419525.e 0000 0001 0690 1414 Saskatchewan Cancer Agency, Saskatoon, Canada
116 https://ror.org/02n415q13 grid.1032.0 0000 0004 0375 4078 Curtin University, Bentley, Australia
117 https://ror.org/009avj582 grid.5288.7 0000 0000 9758 5690 Oregon Health and Science University, Portland, USA
118 https://ror.org/02nkdxk79 grid.224260.0 0000 0004 0458 8737 Virginia Commonwealth University, Richmond, USA
119 https://ror.org/05wvpxv85 grid.429997.8 0000 0004 1936 7531 Tufts University, Medford, USA
120 https://ror.org/00thqtb16 grid.266813.8 0000 0001 0666 4105 University of Nebraska Medical Center, Omaha, USA
121 https://ror.org/01ckdn478 grid.266623.5 0000 0001 2113 1622 University of Louisville, Louisville, USA
122 https://ror.org/02jzgtq86 grid.65499.37 0000 0001 2106 9910 Dana Farber Cancer Institute, Boston, USA
123 https://ror.org/00mkhxb43 grid.131063.6 0000 0001 2168 0066 University of Notre Dame, Notre Dame, USA
124 https://ror.org/0160cpw27 grid.17089.37 University of Alberta, Edmonton, Canada
125 https://ror.org/03dbr7087 grid.17063.33 0000 0001 2157 2938 University of Toronto, Toronto, Canada
126 grid.16563.37 0000000121663741 University of Piemonte Orientale, Vercelli, Italy
127 https://ror.org/05t99sp05 grid.468726.9 0000 0004 0486 2046 University of California, San Francisco, San Francisco, USA
128 grid.25073.33 0000 0004 1936 8227 St. Joseph’s Healthcare Hamilton, and Hamilton Center for Kidney Research, McMaster University, Hamilton, Canada
129 https://ror.org/0168r3w48 grid.266100.3 0000 0001 2107 4242 Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California San Diego, San Diego, USA
130 https://ror.org/030bbe882 grid.11630.35 0000 0001 2165 7640 Universidad de La República, Montevideo, Uruguay
131 https://ror.org/00rcxh774 grid.6190.e 0000 0000 8580 3777 University of Cologne, Cologne, Germany
132 https://ror.org/04dpnfr42 grid.449470.a 0000 0004 0416 6542 Johnson University, Knoxville, USA
133 grid.411377.7 0000 0001 0790 959X Indiana University, Bloomington, USA
134 grid.63054.34 0000 0001 0860 4915 School of Medicine, University of Connecticut, Farmington, USA
135 https://ror.org/05bk57929 grid.11956.3a 0000 0001 2214 904X Stellenbosch University, Stellenbosch, South Africa
136 https://ror.org/02rc97e94 grid.7778.f 0000 0004 1937 0319 University of Calabria, Arcavacata, Italy
137 https://ror.org/05hs6h993 grid.17088.36 0000 0001 2195 6501 Michigan State University, East Lansing, USA
138 https://ror.org/00cvxb145 grid.34477.33 0000 0001 2298 6657 University of Washington, Washington, USA
139 grid.416973.e 0000 0004 0582 4340 Weill Cornell Medicine-Qatar, Ar-Rayyan, Qatar
140 https://ror.org/03ryywt80 grid.256155.0 0000 0004 0647 2973 Gachon University, Seongnam, South Korea
141 https://ror.org/02gz6gg07 grid.65456.34 0000 0001 2110 1845 Florida International University, Miami, USA
142 https://ror.org/05dxps055 grid.20861.3d 0000 0001 0706 8890 California Institute of Technology, Pasadena, USA
143 https://ror.org/00dvg7y05 grid.2515.3 0000 0004 0378 8438 Boston Children’s Hospital, Boston, USA
144 https://ror.org/010x8gc63 grid.25152.31 0000 0001 2154 235X Saskatchewan Cancer Agency and University of Saskatchewan, Saskatchewan, Canada
145 Sino-American Cancer Foundation, Covina, USA
146 https://ror.org/027bh9e22 grid.5132.5 0000 0001 2312 1970 Leiden University, Leiden, The Netherlands
147 grid.430387.b 0000 0004 1936 8796 Rutgers University, Newark, USA
148 grid.417623.5 0000 0004 1758 0566 Core Research Laboratory, ISPRO, Florence, Italy
149 https://ror.org/05dxps055 grid.20861.3d 0000 0001 0706 8890 Caltech, Pasadena, USA
150 University of Alberta, Edmonton, USA
151 grid.482476.b 0000 0000 8995 9090 Montreal Heart Institute and Université de Montréal, Montreal, Canada
152 https://ror.org/01pxwe438 grid.14709.3b 0000 0004 1936 8649 McGill University, Montreal, Canada
153 https://ror.org/008x57b05 grid.5284.b 0000 0001 0790 3681 Antwerp University, Antwerp, Belgium
154 https://ror.org/027ynra39 grid.7644.1 0000 0001 0120 3326 University of Bari Aldo Moro, Bari, Italy
155 https://ror.org/01111rn36 grid.6292.f 0000 0004 1757 1758 Alma Mater Studiorum-University of Bologna, Bologna, Italy
156 https://ror.org/01pbdzh19 grid.267337.4 0000 0001 2184 944X University of Toledo, Toledo, USA
157 https://ror.org/01y3dkx74 grid.419362.b International Institute of Molecular and Cell Biology in Warsaw, Warsaw, Poland
158 https://ror.org/01nrxwf90 grid.4305.2 0000 0004 1936 7988 Infection Medicine, University of Edinburgh The Chancellor’s Building, Edinburgh, UK
159 https://ror.org/03czfpz43 grid.189967.8 0000 0004 1936 7398 Emory University, Atlanta, USA
160 https://ror.org/00b30xv10 grid.25879.31 0000 0004 1936 8972 University of Pennsylvania, Philadelphia, USA
161 https://ror.org/00cfam450 grid.4567.0 0000 0004 0483 2525 Helmholtz Zentrum München, Munich, Germany
162 https://ror.org/03x3g5467 Washington University School of Medicine, St. Louis, USA
163 https://ror.org/030bbe882 grid.11630.35 0000 0001 2165 7640 CENUR Litoral Norte, Universidad de La República, Montevideo, Uruguay
164 https://ror.org/01ryk1543 grid.5491.9 0000 0004 1936 9297 University of Southampton, Southampton, UK
165 grid.1026.5 0000 0000 8994 5086 Centre for Cancer Biology, University of South Australia, Adelaide, Australia
166 https://ror.org/02ymw8z06 grid.134936.a 0000 0001 2162 3504 University of Missouri, Columbia, USA
167 https://ror.org/02pttbw34 grid.39382.33 0000 0001 2160 926X Baylor College of Medicine, Houston, USA
168 https://ror.org/03v76x132 grid.47100.32 0000 0004 1936 8710 Yale University, New Haven, USA
169 https://ror.org/01keh0577 grid.266818.3 0000 0004 1936 914X Reno School of Medicine, University of Nevada, Reno, USA
170 https://ror.org/02gfys938 grid.21613.37 0000 0004 1936 9609 University of Manitoba and CancerCare Manitoba, Winnipeg, Canada
171 https://ror.org/04cvxnb49 grid.7839.5 0000 0004 1936 9721 Goethe University Frankfurt, Frankfurt, Germany
172 grid.413864.c 0000 0004 0420 2582 Oklahoma Medical Research Foundation/Oklahoma City VA Medical Center, Oklahoma City, USA
173 https://ror.org/035z6xf33 grid.274264.1 0000 0000 8527 6890 Oklahoma Medical Research Foundation, Oklahoma City, USA
174 https://ror.org/036rp1748 grid.11899.38 0000 0004 1937 0722 Department of Biomolecular Sciences, School of Pharmaceutical Sciences of Ribeirão Preto, University of São Paulo, Ribeirão Preto, SP Brazil
175 https://ror.org/017cjz748 grid.42687.3f 0000 0004 0381 814X Ulsan National Institute of Science and Technology, Ulsan, South Korea
176 https://ror.org/05kx2e072 0000 0004 0373 6857 University of New Mexico Comprehensive Cancer Center, Albuquerque, USA
177 https://ror.org/017cjz748 grid.42687.3f 0000 0004 0381 814X Ulsan National Institute of Science and Technology (UNIST), Ulsan, South Korea
178 https://ror.org/016476m91 grid.7107.1 0000 0004 1936 7291 University of Aberdeen, Aberdeen, UK
179 https://ror.org/048tbm396 grid.7605.4 0000 0001 2336 6580 University of Turin, Turin, Italy
180 https://ror.org/030bbe882 grid.11630.35 0000 0001 2165 7640 Universidad de La República, CenUR LN, Montevideo, Uruguay
181 grid.428855.6 Navarrabiomed-IdiSNA, Pamplona, Spain
182 Independent, Los Angeles, USA
183 https://ror.org/05dq2gs74 grid.412807.8 0000 0004 1936 9916 Vanderbilt University Medical Center, Nashville, USA
184 https://ror.org/04a7rxb17 grid.18048.35 0000 0000 9951 5557 University of Hyderabad, Hyderabad, India
185 https://ror.org/0384j8v12 grid.1013.3 0000 0004 1936 834X University of Sydney, Sydney, Australia
186 grid.410425.6 0000 0004 0421 8357 City of Hope Medical Center, Duarte, USA
187 https://ror.org/02r109517 grid.471410.7 0000 0001 2179 7643 Weill Cornell Medicine, New York, NY 10065 USA
188 https://ror.org/01pbdzh19 grid.267337.4 0000 0001 2184 944X University of Toledo College of Medicine and Life Sciences, Toledo, USA
189 grid.254444.7 0000 0001 1456 7807 School of Medicine, Wayne State University, Detroit, USA
190 https://ror.org/00te3t702 grid.213876.9 0000 0004 1936 738X University of Georgia, Athens, USA
191 https://ror.org/057q4rt57 grid.42327.30 0000 0004 0473 9646 The Hospital for Sick Children, Toronto, Canada
192 grid.463419.d 0000 0001 0946 3608 United States Department of Agriculture, Agricultural Research Service (USDA-ARS), Washington, DC USA
193 https://ror.org/02fa3aq29 grid.25073.33 0000 0004 1936 8227 McMaster University, Hamilton, Canada
194 https://ror.org/05t99sp05 grid.468726.9 0000 0004 0486 2046 University of California, Riverside, Riverside, USA
195 https://ror.org/047272k79 grid.1012.2 0000 0004 1936 7910 The University of Western Australia, Perth, Australia
196 https://ror.org/02der9h97 grid.63054.34 0000 0001 0860 4915 The University of Connecticut, Storrs, USA
197 grid.26009.3d 0000 0004 1936 7961 Duke University School of Medicine, Durham, USA
198 https://ror.org/043mer456 grid.24434.35 0000 0004 1937 0060 University of Nebraska-Lincoln, Lincoln, USA
199 Sungshin University, Seoul, South Korea
200 https://ror.org/00892tw58 grid.1010.0 0000 0004 1936 7304 University of Adelaide, Adelaide, Australia
201 https://ror.org/03dbr7087 grid.17063.33 0000 0001 2157 2938 University Toronto, Toronto, Canada
202 https://ror.org/05byvp690 grid.267313.2 0000 0000 9482 7121 University of Texas Southwestern Medical Center, Dallas, USA
203 https://ror.org/015w4v032 grid.428469.5 0000 0004 1794 1018 Centro Nacional de Biotecnologia/CSIC, Madrid, Spain
204 grid.5924.a 0000000419370271 Centro de Investigación Médica Aplicada, Pamplona, Spain
205 https://ror.org/02rxc7m23 grid.5924.a 0000 0004 1937 0271 Centro de Investigación Médica Aplicada, Universidad de Navarra, Pamplona, Spain
206 https://ror.org/01y2jtd41 grid.14003.36 0000 0001 2167 3675 University of Wisconsin-Madison, Madison, USA
207 https://ror.org/02ets8c94 0000 0001 2296 1126 Indiana University School of Medicine, Indianapolis, USA
208 https://ror.org/0078xmk34 grid.253613.0 0000 0001 2192 5772 University of Montana, Missoula, USA
209 https://ror.org/00dn4t376 grid.7728.a 0000 0001 0724 6933 Brunel University London, London, UK
210 https://ror.org/0452jzg20 grid.254024.5 0000 0000 9006 1798 Chapman University, Orange, USA
211 grid.416973.e 0000 0004 0582 4340 Weill Cornell Medicine Qatar, Ar-Rayyan, Qatar
212 https://ror.org/002rjbv21 grid.38678.32 0000 0001 2181 0211 Université du Québec À Montréal, Montreal, Canada
213 https://ror.org/05bqach95 grid.19188.39 0000 0004 0546 0241 National Taiwan University, Taipei, Taiwan
214 https://ror.org/049xfwy04 grid.262541.6 0000 0000 9617 4320 Rhodes College, Memphis, USA
215 grid.38142.3c 000000041936754X Harvard School of Public Health, Boston, USA
216 https://ror.org/036nfer12 grid.170430.1 0000 0001 2159 2859 University of Central Florida, Orlando, USA
217 https://ror.org/022kthw22 grid.16416.34 0000 0004 1936 9174 University of Rochester, Rochester, USA
218 https://ror.org/02jqj7156 grid.22448.38 0000 0004 1936 8032 George Mason University, Fairfax, USA
219 https://ror.org/03yj89h83 grid.10858.34 0000 0001 0941 4873 University of Oulu, Oulu, Finland
220 https://ror.org/05n7v5997 grid.476458.c Instituto Investigación Sanitaria La Fe, Valencia, Spain
221 grid.5252.0 0000 0004 1936 973X Ludwig-Maximilians-University, Munich, Germany
222 https://ror.org/05abbep66 grid.253264.4 0000 0004 1936 9473 Brandeis University, Waltham, USA
223 https://ror.org/030bbe882 grid.11630.35 0000 0001 2165 7640 Universidad de La República, CENUR Litoral Norte, Montevideo, Uruguay
224 grid.411424.6 0000 0001 0440 9653 Arabian Gulf University, Manama, Bahrain
225 https://ror.org/013meh722 grid.5335.0 0000 0001 2188 5934 University of Cambridge, Cambridge, UK
226 https://ror.org/0530bdk91 grid.411489.1 0000 0001 2168 2547 University of Magna Graecia, Catanzaro, Italy
227 https://ror.org/002pd6e78 grid.32224.35 0000 0004 0386 9924 Massachusetts General Hospital, Boston, USA
228 https://ror.org/02ymw8z06 grid.134936.a 0000 0001 2162 3504 University of Missouri-Columbia, Columbia, USA
229 https://ror.org/04gr4te78 grid.259670.f 0000 0001 2369 3143 Marquette University, Milwaukee, USA
230 https://ror.org/04a5szx83 grid.266862.e 0000 0004 1936 8163 University of North Dakota, Grand Forks, USA
231 https://ror.org/0213rcc28 grid.61971.38 0000 0004 1936 7494 Simon Fraser University, Burnaby, Canada
232 grid.419404.c 0000 0001 0701 0170 CancerCare Manitoba Research Institute (CCMR), Winnipeg, Canada
233 https://ror.org/00rqy9422 grid.1003.2 0000 0000 9320 7537 The University of Queensland, Brisbane, Australia
234 https://ror.org/005cmms77 grid.419404.c 0000 0001 0701 0170 University of Manitoba and CancerCare Manitoba Research Institute, Winnipeg, Canada
235 grid.479509.6 0000 0001 0163 8573 Sanford Burnham Prebys, La Jolla, USA
236 https://ror.org/04dkp9463 grid.7177.6 0000 0000 8499 2262 University of Amsterdam, Amsterdam, The Netherlands
237 grid.208078.5 0000000419370394 UConn Health, Farmington, USA
238 https://ror.org/04twxam07 grid.240145.6 0000 0001 2291 4776 The University of Texas MD Anderson Cancer Center, Houston, USA
239 https://ror.org/02svzjn28 grid.412870.8 0000 0001 0447 7939 Walter Sisulu University, Mthatha, South Africa
240 https://ror.org/02smfhw86 grid.438526.e 0000 0001 0694 4940 Virginia Tech, Blacksburg, USA
241 https://ror.org/048sx0r50 grid.266436.3 0000 0004 1569 9707 University of Houston, Houston, USA
242 https://ror.org/05vt9qd57 grid.430387.b 0000 0004 1936 8796 Rutgers University, New Brunswick, USA
243 https://ror.org/01keh0577 grid.266818.3 0000 0004 1936 914X University of Nevada, Reno, USA
244 https://ror.org/02tzt0b78 grid.4807.b 0000 0001 2187 3167 Universidad de León, León, Spain
245 grid.67105.35 0000 0001 2164 3847 School of Medicine, Case Western Reserve University, Cleveland, USA
246 grid.208078.5 0000000419370394 School of Medicine, UConn Health, Farmington, USA
247 grid.412750.5 0000 0004 1936 9166 University of Rochester Medical Center, Rochester, USA
248 https://ror.org/0452jzg20 grid.254024.5 0000 0000 9006 1798 Chapman University School of Pharmacy, Irvine, USA
249 https://ror.org/02gdzyx04 grid.267457.5 0000 0001 1703 4731 University of Winnipeg/St. Boniface Research Centre, Winnipeg, Canada
250 https://ror.org/007ps6h72 grid.270240.3 0000 0001 2180 1622 Fred Hutchinson Cancer Center, Seattle, USA
251 https://ror.org/042dsac10 grid.461899.b Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Saarbrücken, Germany
252 https://ror.org/04qtj9h94 grid.5170.3 0000 0001 2181 8870 Technical University of Denmark, Kongens Lyngby, Denmark
253 https://ror.org/00wmhkr98 grid.254250.4 0000 0001 2264 7145 The City College of New York, New York, USA
254 grid.468147.8 Children’s Cancer, Therapy Development Institute (Cc-TDI), Beaverton, USA
255 https://ror.org/003qdfg20 grid.412664.3 0000 0001 0284 0976 Soka University, Hachioji, Japan
256 grid.5326.2 0000 0001 1940 4177 Institute of Clinical Physiology, National Research Council, Pisa, Italy
257 https://ror.org/00wfvh315 grid.1037.5 0000 0004 0368 0777 Charles Sturt University, Bathurst, Australia
258 grid.4367.6 0000 0001 2355 7002 Washington University, St Louis, USA
259 https://ror.org/03h2bxq36 grid.8241.f 0000 0004 0397 2876 University of Dundee, Dundee, UK
260 https://ror.org/03vs03g62 grid.482476.b 0000 0000 8995 9090 Montreal Heart Institute, Montreal, Canada
261 https://ror.org/00hj8s172 grid.21729.3f 0000 0004 1936 8729 Columbia University Vagelos College of Physicians and Surgeons, Columbia, USA
262 https://ror.org/05vzafd60 grid.213910.8 0000 0001 1955 1644 Georgetown University, Washington, USA
263 https://ror.org/00py81415 grid.26009.3d 0000 0004 1936 7961 Duke University, Durham, USA
264 https://ror.org/02rxc7m23 grid.5924.a 0000 0004 1937 0271 Center for Applied Medical Research, University of Navarra, Pamplona, Spain
265 https://ror.org/0220mzb33 grid.13097.3c 0000 0001 2322 6764 King’s College London, London, UK
266 https://ror.org/00dvg7y05 grid.2515.3 0000 0004 0378 8438 Precision Vaccines Program, Division of Infectious Diseases, Boston Children’s Hospital, Boston, USA
267 grid.270240.3 0000 0001 2180 1622 Fred Hutchinson Cancer Research Center, Seattle, USA
268 https://ror.org/05ect4e57 grid.64337.35 0000 0001 0662 7451 Louisiana State University, Baton Rouge, USA
269 https://ror.org/052czxv31 grid.148374.d 0000 0001 0696 9806 Massey University, Palmerston North, New Zealand
270 https://ror.org/0207ad724 grid.241167.7 0000 0001 2185 3318 Wake Forest University School of Medicine, Winston-Salem, USA
271 https://ror.org/00f1zfq44 grid.216417.7 0000 0001 0379 7164 Central South University, Changsha, China
272 https://ror.org/040kfrw16 grid.411023.5 0000 0000 9159 4457 SUNY Upstate Medical University, Syracuse, USA
273 https://ror.org/052gg0110 grid.4991.5 0000 0004 1936 8948 University of Oxford, Oxford, UK
274 grid.7839.5 0000 0004 1936 9721 Goethe-University, Frankfurt, Frankfurt, Germany
275 https://ror.org/05591te55 grid.5252.0 0000 0004 1936 973X Institute for Cardiovascular Prevention (IPEK), Ludwig-Maximilians-Universität München, Munich, Germany
276 grid.38142.3c 000000041936754X Harvard T.H. Chan School of Public Health, Boston, USA
277 grid.189504.1 0000 0004 1936 7558 School of Medicine, Boston University, Boston, USA
278 https://ror.org/02f6dcw23 grid.267309.9 0000 0001 0629 5880 University of Texas Health Science Center at San Antonio, San Antonio, USA
279 https://ror.org/040af2s02 grid.7737.4 0000 0004 0410 2071 University of Helsinki, Helsinki, Finland
280 https://ror.org/050kf9c55 grid.465543.5 0000 0004 0435 9002 Wadsworth Center, NYSDOH, Albany, USA
281 https://ror.org/0384j8v12 grid.1013.3 0000 0004 1936 834X The University of Sydney, Sydney, Australia
2 4 2024
2 4 2024
2024
14 752615 9 2023
15 2 2024
© The Author(s) 2024, corrected publication 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
High throughput screening (HTS) is routinely used to identify bioactive small molecules. This requires physical compounds, which limits coverage of accessible chemical space. Computational approaches combined with vast on-demand chemical libraries can access far greater chemical space, provided that the predictive accuracy is sufficient to identify useful molecules. Through the largest and most diverse virtual HTS campaign reported to date, comprising 318 individual projects, we demonstrate that our AtomNet® convolutional neural network successfully finds novel hits across every major therapeutic area and protein class. We address historical limitations of computational screening by demonstrating success for target proteins without known binders, high-quality X-ray crystal structures, or manual cherry-picking of compounds. We show that the molecules selected by the AtomNet® model are novel drug-like scaffolds rather than minor modifications to known bioactive compounds. Our empirical results suggest that computational methods can substantially replace HTS as the first step of small-molecule drug discovery.

Subject terms

Drug discovery
High-throughput screening
Virtual screening
Machine learning
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Despite present interest in AI/ML and thirty years of case studies1–4, computational screening techniques have achieved limited adoption within the pharmaceutical industry. A recent investigation into the origins of 156 clinical candidates5 found that only 1% came from virtual screening; in contrast, over 90% of clinical candidates were derived from patent busting or high throughput screening (HTS). Unfortunately, these sources are increasingly challenged, given the pharmaceutical industry’s shift to novel target classes, such as proximity-induced protein degradation6, protein–protein interactions7, and RNA targeting8.

Currently, HTS is the critical tool in drug discovery, providing most novel scaffolds of recent clinical candidates5,9,10. These initial starting points crucially shape the course of downstream medicinal chemistry efforts, as most drugs preserve at least 80% of the scaffold of the initially identified lead11. Despite these foundational contributions, HTS suffers from practical limitations. Principally, HTS, like all physical experiments, requires that the compounds exist. However, with the advent of synthesis-on-demand libraries, most commercially-available molecules have yet to be synthesized. Still, they can be made and delivered for testing in a matter of weeks12–14. These libraries comprise trillions of molecules14,15 that exemplify millions of otherwise-unavailable scaffolds12, providing an opportunity to substantially expand the scope and diversity of available chemical space explored in the standard drug discovery process.

Computational approaches unlock this opportunity by reversing the requirement to make molecules before testing them. When computational experiments replace HTS as the primary screen, molecules are tested before they are made, and the results from these experiments can inform which molecules are worth synthesizing. Computational experiments further promise to improve upon HTS in terms of cost, speed, need to produce significant quantities of protein16, effort of miniaturizing assay formats while maintaining experimental integrity17–19, and reducing false-positive and false-negative rates16,20–23 including artifacts from aggregation, covalent modification of the target, autofluorescence, or interactions with the reporter rather than the target20,24,25. Historical computational techniques such as ligand-based QSAR26–28, structure-based docking29,30, and machine learning31,32 purport to address these limitations of physical screening methods. Unfortunately, these techniques have not replaced HTS; in fact, despite increasing interest in ML, the proportion of drugs discovered with computational techniques has remained steady over the past decades5,10.

Because there will always be individual targets for which one screening technique can identify more hits than another, the key question governing if computation is ready to be the default hit discovery technique is whether computational screens can identify hits successfully across a broad range of diverse targets. Unfortunately, despite excellent benchmark accuracies33–35, prospective discovery accuracy remains modest33,36,37. For example, Cerón-Carrasco38 reported over 700 virtual screens against the SARS-CoV-2 main protease. However, when the author sought to validate the computational predictions via physical experiments, the identified compounds were barely active (800uM). Computational approaches have also been limited by a need for extensive target-specific training data31,39–41, a requirement for high-quality X-ray crystal structures42,43, dependence on human adjudication (so-called ‘cherry-picking’)12, or a limited domain of applicability44–48. Even recent systems have demonstrated utility only in identifying minor variants of known molecules for well-studied proteins with tens of thousands of known binders in their training data49,50. Figure 1 exemplifies the striking similarities between recently ML-developed compounds and their preceding published chemical matter. This is particularly concerning, as a myopic focus on well-studied proteins has been identified as a cause of low productivity in pharmaceutical discovery51.Figure 1 Pairs of representative compounds extracted from AI patents (right) and corresponding prior patents (left) for clinical-stage programs (CDK792,93, A2Ar-antagonist94,95, MALT196,97, QPCTL98,99, USP1100,101, and 3CLpro102,103). The identical atoms between the chemical structures are highlighted in red.

Nevertheless, we have observed that deep learning approaches are not as limited as these historical examples would imply. Using our AtomNet52–54 screening system, we have previously reported success in finding novel scaffolds for targets without known ligands55–57, X-ray crystal structures56–60, or both56,57, as well as challenging modulation via protein–protein interaction59,61 or allosteric binding60 (see Supplementary Table S1 for examples). However, individual examples do not demonstrate the overall success of such deep learning systems. We therefore report our internal discovery efforts against 22 targets of pharmaceutical interest. We then attempted to further assess the generalizability and robustness of deep learning predictive systems by identifying bioactive molecules for a diverse set of targets. We partnered with 482 academic labs and screening centers, from 257 different academic institutions across 30 countries, through our academic collaboration program, the Artificial Intelligence Molecular Screen (AIMS). This collaboration afforded an opportunity to prospectively evaluate the utility of the AtomNet model as a primary screen across a broad range of diverse, challenging, and realistic targets. In aggregate, we report successes and failures from 318 prospective experiments and evaluate our AtomNet machine-learning technology’s ability to serve as a viable alternative to physical HTS campaigns.

Results

We investigated the ability of deep learning-based methods to identify novel bioactive chemotypes by applying the AtomNet model to identify hits for 22 internal targets of pharmaceutical interest. We also explored the breadth of applicability of this approach by attempting to identify drug-like hits in single-dose screens for 296 academic targets, of which 49 were followed up with dose–response experiments, and 21 were further validated by exploring analogs of the initial hits. The average hit rate for our internal projects (6.7%) was comparable to the hit rate for our academic collaborations (7.6%).

Internal portfolio validation

As part of Atomwise’s internal drug discovery efforts, we used the AtomNet model instead of high-throughput or DNA-encoded library (DEL) screening. We screened a 16-billion synthesis-on-demand chemical space62, which is several thousand times larger than HTS libraries and even exceeds the size of most DELs without suffering limitations of DNA-compatible chemistry16,23. Each screen requires over 40,000 CPUs, 3,500 GPUs, 150 TB of main memory, and 55 TB of data transfers. We describe the protocol in detail in the Methods section; briefly, we computationally scored each catalog compound after removing molecules that were prone to interfere with the assays or were too similar to known binders of the target or its homologs. The neural network analyzes and scores the 3D coordinates of each generated protein–ligand co-complex, producing a list of ligands ranked by their predicted binding probability. Our workflow then clusters the top-ranked molecules to ensure diversity and algorithmically selects the highest-scoring exemplars from each cluster. At no point are compounds manually cherry-picked. The molecules were synthesized at Enamine (https://enamine.net) and quality controlled by LC–MS to purity > 90%, in agreement with HTS standards63. Hits were further validated using NMR. We then physically tested, on average, 440 compounds per target at reputable contract research organizations (CROs), while attempting to mitigate assay interferences such as aggregation and oxidation with standard additives (e.g., Tween-20, Triton-X 100, and dithiothreitol (DTT)). We describe the assay protocols in detail in the Supplementary Data S1.

We describe the results of the 22 experiments in Table 1. In 91% of the experiments, we identified single-dose (SD) hits that were reconfirmed in dose–response (DR) experiments. The average target DR hit rate was 6.7% compared to 8.8% from the SD screens. Only 16 of the 22 projects were structurally enabled with X-ray crystallography; one used a cryo-EM structure, while five used homology models with an average sequence identity of 42% to their template protein. The DR hit rate for the cryo-EM project was 10.56%, while the average hit rate for the homology models was a similar 10.8%.Table 1 Results from 22 Atomwise internal programs.

Gene name	# of compounds tested	SD hit rate (%)	DR hit rate (%)	Potency range (IC50/Ki, uM)	# of analog tested	SD analog hit rate (%)	DR analog hit rate (%)	Analog potency range (IC50/Ki, uM)	
ASAH1	376	10.64	7.71	0.3–102	–	–	–	–	
AXL	597	12.06	8.21	0.181–71	3200	35.59	33.56	0.079–86	
BCL2	422	3.08	0.00	–	–	–	–	–	
CBLB	422	1.66	0.00	–	–	–	–	–	
CDK5	786	10.69	10.43	0.049–79	587	47.53	43.61	0.43–76	
CDK7	786	10.69	10.56	0.099–60	735	28.44	27.35	0.191–10	
GFPT1	384	6.51	2.34	31–86	734	24.93	24.11	1–194	
KCNT1	416	9.62	7.69	1.1–30	–	–	–	–	
KDM6A	356	3.93	1.12	24–58	–	–	–	–	
LATS1	418	18.18	17.94	0.077–82	841	51.72	45.78	0.034–98	
MC2R	208	11.54	9.62	16–68	419	39.38	38.42	2.4–97	
MDM4	422	2.37	0.47	5.9–29.8	192	18.23	18.23	4.4–90	
NT5E	335	1.49	0.30	176	221	9.95	1.81	8.3–65	
PARG	334	7.78	7.78	15–250	–	–	–	–	
PARP14	576	5.38	2.95	3–96	616	26.46	26.30	0.2–95	
POLQ	330	11.82	11.52	1.2–49	559	11.27	8.77	1.5–42	
PPARA	422	4.03	0.24	131	211	14.22	3.79	59–95	
PPM1D	530	11.89	6.98	4.5–98	–	–	–	–	
PRMT5	422	4.03	0.95	7.2–79	415	7.95	5.54	19–114	
PRODH2	542	2.77	1.11	15–84	–	–	–	–	
TYK2	189	38.10	34.39	0.016–9	457	71.33	60.39	0.006–10	
VCP	416	4.81	4.81	2.4–64	738	–	–	–	
SD and DR denote single-dose and dose–response, respectively.

We then advanced 14 projects with at least one dose-responsive scaffold to a round of analog expansion. We found new bioactive analogs in the SD screen for all projects, with an average hit rate of 29.8%. Further validation with DR resulted in an average hit rate of 26% per project, which compares favorably with typical HTS hit rates ranging from 0.151 to 0.001%64,65. We note that the size and chemical diversity within and between physical66 and virtual14 HTS libraries prevent an explicit evaluation of the methods over the same chemical space. The most potent analogs ranged from single-digit nanomolar, against a kinase, to double-digit micromolar, against a transcription factor (Supplementary Table S2). Additionally, we present two internal studies in detail. For Large Tumor Suppressor Kinase 1 (LATS1), we identified potent compounds despite the lack of a crystal structure or known active compounds. For ATP-driven chaperone Valosin Containing Protein (VCP) we identified novel allosteric and orthosteric modulators.

Academic validation

In addition to our internal discovery efforts, we performed virtual screens for 296 targets, comprising more than 20 billion individual neural network scores of generated protein–ligand co-complexes. We purchased, on average, 85 off-the-shelf commercially available compounds, quality controlled by NMR and LC–MS to > 90% purity63, and plated in a single 96-well plate. The compounds were then physically screened for activity against the target of interest in single-dose assays (see Supplemental Data S1 for assay protocols). As with HTS primary screens, additional characterization studies are required to validate the initially identified hits so, in 49 projects, we performed dose–response studies and analog expansion. We present a summary of our results in Supplementary Table S3.

Figure 2 illustrates the distributions of projects across therapeutic areas, protein families, and assay types. Every major therapeutic area is represented, with the most frequent area being oncology, comprising 35% of projects, followed by infectious diseases and neurology, comprising 27% and 9% of projects, respectively. Breaking down the projects by protein families reveals that all major enzyme classes are represented, with enzymes comprising 59% of the targets and membrane proteins such as GPCR, transporters, and ion channels, representing 12% of the targets. Working on a large and diverse set of therapeutic targets requires a heterogeneous collection of biological assays; 20% of the assays measured direct binding, whereas 56% and 20% were functional and phenotypic.Figure 2 The distributions of 296 AIMS projects across assay types used in the primary screen, research areas, target classes, and further breakdown to enzyme classes when applicable.

In 215 projects, we identified at least one bioactive compound for the target in a biochemical or cell-based assay. This 73% success rate substantially improves over the ∼50% success rate for HTS21,67. On average, we screened 85 compounds per project and discovered 4.6 active hits, with an average hit rate of 5.5%. For the subset of targets where we found any hits, the average was 6.4 hits per project. Thus, we achieved an average hit rate of 7.6%, which again compares favorably with typical HTS hit rates. See Supplementary Material S1 for all assay definitions and conditions. Supplementary Table S4 shows a representative bioactive compound from each of the 215 successful projects, and Supplementary Fig. S2 shows that the physicochemical properties of the identified hits are largely druglike and Lipinski-compliant.

The AtomNet technology robustly identified active molecules, even for targets that lacked prior on-target bioactivity data. This ability to identify hits for previously undrugged targets is critical if machine learning-based approaches are to replace HTS as the default primary screening approach. For 207 out of the 296 targets (70%), the training data available for AtomNet models lacked a single active molecule for that target or any closely related protein (i.e., proteins with sequence identity greater than 70%). We interpret this as evidence of the ability of properly-architected machine learning systems to extrapolate to novel biological space. Figure 3A illustrates the hit rate versus the number of training examples available to our model. Although previous computational approaches typically require thousands of on-target training examples31,39,42, the lack of correlation between training examples and hit rate (R2 = 0.0021, p-value = 0.43) shows that our ML algorithm is agnostic to the availability of such data. We achieved an average success rate of 75% and hit rates of 5.3% when no training data was available, comparable to the 67% and 6.1% success and hit rates achieved when binding data was available in the training set. Interestingly, we also do not see a significant increase in hit rate attributable to the proportion of binding data available for a target (R2 = 0.008, p-value = 0.39). This reflects the robustness of the screening protocol and the chemical dissimilarity of scaffolds identified by AtomNet models to previously known bioactive compounds.Figure 3 (A) An illustration of the hit rate versus the number of training examples available to our model. Each point represents a project, with the x-axis denoting the number of active molecules in our training for the target protein or homologs and the y-axis denoting the hit rate of the project (the percentage of molecules tested in the project that were active). The model shows no dependence on the availability of on-target training examples. For 70% of the targets, the AtomNet model training data lacked any active molecules for that target or any similar targets with greater than 70% sequence identity, yet the model achieved a hit rate of 5.3% compared to 6.1% when on-target data was available. (B) The distribution of similarities between hits and their most-similar bioactive compounds in our training data. Our screening protocol ensures that the compounds subjected to physical testing are not similar to known active compounds or close homologs (< 0.5 Tanimoto similarity using ECFP4, 1024 bits). Because 70% of the AIMS targets had no annotated bioactivities in our training dataset, hits identified in these projects have a similarity value of zero.

Next, we assessed the ability of the AtomNet models to identify novel scaffolds. This is a critical capability for primary screens, as follow-up assays tend to work within the chemical space uncovered in the initial screen. The task of novel scaffold identification appears in two distinct scenarios: (1) when no scaffold is known for the target and we wish to identify the first scaffold, and (2) when some scaffolds are known but we wish to identify dissimilar scaffolds because novel chemical matter can yield improved selectivity, toxicity, pharmacokinetics, or patentability. Performance of AtomNet models for the first scenario, when no scaffolds for the target existed in the AtomNet model training data, was evaluated on 70% of the targets, where the training data contained no active molecules for the target or its homologs (vide supra). We achieved an average hit rate of 5.3% for targets with no training data. For the second scenario, we analyzed the similarity of the identified hits to known bioactive compounds in our training data (Fig. 3B). Our screening protocol ensures that the compounds subjected to physical testing are not similar to known active compounds or close homologs (< 0.5 Tanimoto similarity using ECFP468, 1024 bits). We interpret this as evidence of the ability of properly-architected machine learning systems to extrapolate to novel chemical space as well. For cases where training data was available (i.e., the Tanimoto similarity is above zero), the similarity distribution is close to the one expected by random compound pairs69. The novelty of the small-molecule structures is striking because target-specific machine-learning algorithms tend to uncover highly similar analogs for known bioactive molecules50,70,71. The superior performance of the AtomNet model is expected, considering the bias-variance tradeoff72 in machine learning algorithms. Because the AtomNet convolutional neural network is a global model, concurrently trained on millions of bioactivities, hundreds of thousands of small molecules, and thousands of protein binding sites, it can reduce both bias and variance of the model compared to target-specific ones33. Specifically, our global model can benefit from multiple levels of information captured in the structures of the small molecules, the sequences of the target proteins, and the three-dimensional interactions between the two.

AtomNet also successfully identified active molecules when there was no X-ray crystal structure of the receptor. Figure 4A compares the hit rates obtained with 3-dimensional crystal structures, cryo-EM, and homology modeling. We did not attempt to select targets based on the similarity to the template but rather used the best template available. We observe no substantial difference in success rate between the three, in contrast to the common challenges in using homology models or low-precision structures for structure-based discovery42,43,73. We achieved average hit rates of 5.6%, 5.5%, and 5.1% for crystal structures, cryo-EM, and homology modeling. We also successfully identified active compounds in projects with NMR structures, but the number of such targets is too small to make statistically-robust claims.Figure 4 Hit rates obtained for the 296 AIMS projects. (A) A comparison of hit rates using X-ray crystallography, NMR, Cryo-EM, and homology for modeling the structure of the proteins. Each point represents a project with the x-axis denoting the hit rate of the project (the percentage of molecules tested in the project that were active). The number of projects of each type is given in parentheses. We observed no substantial difference in success rate between the physical and the computationally inferred models. We achieved average hit rates of 5.6%, 5.5%, and 5.1% for crystal structures, cryo-EM, and homology modeling, respectively. The number of projects using NMR structures is too small to make statistically-robust claims. (B) A comparison of hit rates observed for traditionally challenging target classes such as protein–protein interactions (PPI) and allosteric binding. Of the 296 projects, 72 targeted PPIs and 58 allosteric binding sites. The average hit rates were 6.4% and 5.8% for PPIs and allosteric binding, respectively. (C) Comparison of hit rates observed for different target classes and (D) enzyme classes. No protein or enzyme class falls outside the domain of applicability of the algorithm.

An interesting demonstration of the robustness of the AtomNet model to low data and poorly characterized protein structure is its ability to identify novel hits for traditionally challenging target classes such as protein–protein interaction (PPI) sites and allosteric binding sites (Fig. 3B). Of the 296 projects, 72 targeted PPIs and 58 allosteric binding sites. We identified hits for 53 (74%) PPI sites and 46 (79%) allosteric sites, with 13 projects representing allosteric sites at PPI interfaces. The average hit rate was 6.4% and 5.8% for PPIs and allosteric binding sites, respectively. The algorithm's success in these target classes, which often suffer from poorly characterized binding sites and a lack of bioactivity training data, is not surprising because Fig. 2A shows that our model is largely not dependent on the availability of on-target training data.

Finally, we investigated whether the algorithm exhibits domain of applicability limitations regarding different protein classes. Figures 4C and 3D illustrate the hit rate observed for each protein and enzyme class. No protein or enzyme class falls outside the domain of applicability of the algorithm, demonstrating that machine learning-based approaches are well-suited as a default technology for new scaffold identification. The hit rate for nuclear receptors is an outlier, with seemingly better accuracy than other classes, but a single data point is not statistically meaningful.

Dose–response validation studies

We performed additional validation studies for 49 AIMS projects with at least one reported hit. The objective of the validation studies was to establish dose–response (DR) relationships for the single-dose (SD) hits. We describe the protocol of the DR experiments in the Methods section. Briefly, we performed dose–response measurements for the reported hits from the single-dose primary screens. DR was determined using the same assay and screening protocol as the single-dose screens, at the same lab, and with the same personnel. Full dose response curves were obtained in most cases, however in some instances a full curve was not obtained, or concentration dependent activity was qualitatively determined by testing at concentrations other than that for the primary screen. The distribution of assay types and target classes for the projects selected for DR validation also was similar to that of the AIMS projects (Supplementary Fig. S3).

We describe the results of the DR experiments in Supplementary Table S5. In 84% of the experiments, we validated at least one SD hit and got a DR readout. The median activity for the total of 144 DR measurements was 15.4 µM (which compares favorably with HTS25,74), of which 13% showed sub-µM potency. Overall, we achieved an average of 2.8 hits per validation study, resulting in a hit rate of 51%. The false positive rate of 49% observed in these experiments is favorably compared to HTS’ which can be as high as 95%20,75. This difference in false positive rates may stem from the comparative ease and robustness of the low-throughput assay format we employed versus high-throughput assay. Representative dose–response curves for each of the 49 projects are shown in Supplementary Table S6.

Analog validation studies

For a subset of 21 projects, we further validated hits with DR activity by testing analogs of the active compounds. In those cases, we used the AtomNet platform to search a purchasable space for additional bioactive compounds chemically analogous to the SD hits. We selected up to 35 additional compounds for testing, including the active compounds from the SD screens.

We describe the results of the analoging experiments in Supplementary Table S7. We identified additional analogs with DR readouts for 16 projects (76%). The median DR activity of the 154 validated analogs was 7.4 µM compared to the median of 15.4 µM of the parent compound (Supplementary Fig. S4).

Methods

Screening protocols

AIMS screening protocol

We began by evaluating screening libraries of millions of catalog compounds from commercial vendors MCule (10 M)76 and Enamine in-stock (2.5 M)77. We then selected a drug-like subset via algorithmic filtering by applying Eli Lilly medicinal chemistry filters78 and removing likely false positives, such as aggregators, autofluorescers, and PAINS79 (see Fig. 2 for the distributions of drug-like properties of the SD hits). The resulting library was virtually screened against the target of interest, removing any molecules with greater than 0.5 Tanimoto similarity in ECFP4 space to any known binders of the target and its homologs within 70% sequence identity. For kinase targets, we extend the exclusion to the whole kinome. The binding site was defined using co-complexes, mutagenesis studies, co-complexes of homologs, or by identifying potential sites using ICM Pocket Finder80 or Fpocket81. Some were orthosteric, while others were allosteric, or as yet unestablished biological functions. In 64 cases, we built homology models using the closest sequence, with an average sequence similarity of 54%. We clustered the top 30,000 molecules using the Butina82 algorithm with a Tanimoto similarity cutoff of 0.35 in ECFP4 space, selecting the highest-scoring exemplars. Additional computed physico-chemical property filters were applied as needed. At no point were compounds cherry-picked. We purchased, on average, 85 compounds, quality controlled by LC–MS to > 90% purity, generally dispensed as 10 mM DMSO stocks plated in a single 96-well plate. In addition, two vials of DMSO-only negative controls were included before scrambling the compound locations on the plate, by the supplier, for blinded experimental testing. To further control for potential artifacts, we removed compounds that showed measurable activity toward more than one target from the analysis.

Dose–response and analoging validation screening protocol

We considered advancing AIMS projects to additional validation studies based on the ability to reorder at least some of the initial SD hits, the availability of chemical analogs in the screening library to the initial hits, the capability to perform dose–response experiments, and the ability of the collaborators to perform additional screens and return results promptly.

We performed two sets of experiments: DR validation of the SD hits from AIMS and analoging with DR readouts. We performed DR measurements using the same assays and protocols as SD.

We performed an analoging round by identifying, for each AIMS hit, its 1000 nearest neighbors from the Mcule library76, using molecular fingerprints similarity68. We augmented the set with additional analogs using substructure83 or FTrees84 searches, if needed. We used an AtomNet regression model, trained to predict quantitative bioactivities (e.g., IC50 or Ki), to score and rank the analogs. A set of 20—35 compounds from the analogs space of an initial hit were then obtained based on similarity and top scores from the AtomNet model for testing.

Internal portfolio screening protocol

We followed a protocol similar to the AIMS screen with a few deviations. First, we used the Enamine REAL library of over 16 billion compounds62. Second, we used an ensemble of six AtomNet models for the screens. Last, on average, we selected a set of 440 compounds for testing.

The analoging protocol is similar to the AIMS validation studies, with the following deviations. First, we used the Enamine REAL library for analog search. Second, we selected an average of 676 analogs per project. Third, the analog search protocol was more complex, pulling nearest neighbors based on maximum common substructure and graph edit distance in addition to the ECFP4-based one.

AtomNet® model architecture

We previously published in detail52,53,55,58,59,61,85,86 during the course of the AIMS program, and we described the most recent version of the AtomNet model architecture in detail elsewhere53. We provide a brief description below.

The AtomNet model is a Graph Convolution Network architecture with atoms represented as vertices and pair-wise, distance-dependent, edges representing atom proximities. The input is a graph network of features characterizing the atom types and topologies of an ensemble of protein–ligand complexes. Receptor atoms more than 7 Å away from any ligand atom are excluded from the complexes, and each node in the graph is associated with a feature vector representing the atom type using Sybyl typing87.

The network has five graph convolutional blocks. In the first two graph convolution blocks, all ligand and receptor atoms 5 Å apart from each other are considered, and 64 filters per block are used. In the third block, the cutoff radius and filters are increased to 7 Å and 128, respectively. Only ligand features in the last two blocks are considered without changing the threshold cutoff or the number of filters. Finally, the sum-pool of the ligand-only layer creates a 3-task layer on top of the network. That multi-task layer predicts three endpoints: bioactivity, pose quality, and a physics-based docking score88.

We trained an ensemble of 6 models, splitting the training data into sixfold cross-validation sets based on a protein sequence similarity cutoff of 70%. Then, each model in the ensemble was trained on a different fold for 10 epochs, using the ADAM optimizer89 with a learning rate of 0.001, and targets were sampled with replacement, proportional to the number of active compounds associated with that target.

Data

All data generated or analyzed during this study are included in this published article (and its supplementary information S1 files). Boxplots illustrations show the quartiles (Q1 and Q3) of the dataset while the whiskers extend to show the rest of the distribution, except for points that are determined to be “outliers” (1.5 × of the inter-quartile range, as implemented in the Seaborn and Matplotlib toolboxes90,91).

Conclusion

HTS is the most widely-used tool for hit discovery for new targets. Unfortunately, all physical screening methods share the critical limitation that a molecule must exist to be screened. Computational methods enable a fundamental shift to a test-then-make paradigm. In this work, we report on 318 projects (22 internal projects and 296 collaborations) where we used the AtomNet platform as the primary screening tool coupled with low-throughput physical screens as validation. The AtomNet technology can identify bioactive scaffolds across a wide range of proteins, even without known binders, X-ray structures, or manual cherry-picking of compounds. Our empirical results suggest that machine learning approaches have reached a computational accuracy that can replace HTS as the first step of small-molecule drug discovery.

Supplementary Information

Supplementary Information 1.

Supplementary Information 2.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-54655-z.

Acknowledgements

See Supplementary section S2.

Author contributions

All authors have contributed to the publication, being variously involved in technology development, experimental protocol designs, experimental performance, data acquisition, statistical analysis, and manuscript writing.

Data availability

All data generated or analyzed during this study are included in this published article and its supplementary information files.

Competing interests

The authors affiliated with Atomwise declare the existence of a financial competing interest.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

A list of authors and their affiliations appears at the end of the paper.
==== Refs
References

1. Kuntz ID Structure-based strategies for drug design and discovery Science 1992 257 1078 1082 10.1126/science.257.5073.1078 1509259
Kuntz, I. D. Structure-based strategies for drug design and discovery. Science 257, 1078–1082 (1992).1509259 10.1126/science.257.5073.1078
2. Bajorath J Integration of virtual and high-throughput screening Nat. Rev. Drug Discov. 2002 1 882 894 10.1038/nrd941 12415248
Bajorath, J. Integration of virtual and high-throughput screening. Nat. Rev. Drug Discov. 1, 882–894 (2002).12415248 10.1038/nrd941
3. Walters WP Stahl MT Murcko MA Virtual screening—an overview Drug Discov. Today 1998 3 160 178 10.1016/S1359-6446(97)01163-X
Walters, W. P., Stahl, M. T. & Murcko, M. A. Virtual screening—an overview. Drug Discov. Today 3, 160–178 (1998).10.1016/S1359-6446(97)01163-X
4. Ring CS Structure-based inhibitor design by using protein models for the development of antiparasitic agents Proc. Natl. Acad. Sci. USA. 1993 90 3583 3587 10.1073/pnas.90.8.3583 8475107
Ring, C. S. et al. Structure-based inhibitor design by using protein models for the development of antiparasitic agents. Proc. Natl. Acad. Sci. USA. 90, 3583–3587 (1993).8475107 10.1073/pnas.90.8.3583
5. Brown DG An analysis of successful hit-to-clinical candidate pairs J. Med. Chem. 2023 10.1021/acs.jmedchem.3c00521 37996079
Brown, D. G. An analysis of successful hit-to-clinical candidate pairs. J. Med. Chem.10.1021/acs.jmedchem.3c00521 (2023).37996079 10.1021/acs.jmedchem.3c00521
6. Békés M Langley DR Crews CM PROTAC targeted protein degraders: The past is prologue Nat. Rev. Drug Discov. 2022 21 181 200 10.1038/s41573-021-00371-6 35042991
Békés, M., Langley, D. R. & Crews, C. M. PROTAC targeted protein degraders: The past is prologue. Nat. Rev. Drug Discov. 21, 181–200 (2022).35042991 10.1038/s41573-021-00371-6
7. Lu H Recent advances in the development of protein–protein interactions modulators: Mechanisms and clinical trials Signal Transduct. Target. Ther. 2020 5 1 23 32296011
Lu, H. et al. Recent advances in the development of protein–protein interactions modulators: Mechanisms and clinical trials. Signal Transduct. Target. Ther. 5, 1–23 (2020).32296011
8. Childs-Disney JL Targeting RNA structures with small molecules Nat. Rev. Drug Discov. 2022 21 736 762 10.1038/s41573-022-00521-4 35941229
Childs-Disney, J. L. et al. Targeting RNA structures with small molecules. Nat. Rev. Drug Discov. 21, 736–762 (2022).35941229 10.1038/s41573-022-00521-4
9. Brown DG Boström J Where do recent small molecule clinical development candidates come from? J. Med. Chem. 2018 61 9442 9468 10.1021/acs.jmedchem.8b00675 29920198
Brown, D. G. & Boström, J. Where do recent small molecule clinical development candidates come from?. J. Med. Chem. 61, 9442–9468 (2018).29920198 10.1021/acs.jmedchem.8b00675
10. Dragovich PS Haap W Mulvihill MM Plancher J-M Stepan AF Small-molecule lead-finding trends across the roche and genentech research organizations J. Med. Chem. 2022 65 3606 3615 10.1021/acs.jmedchem.1c02106 35138850
Dragovich, P. S., Haap, W., Mulvihill, M. M., Plancher, J.-M. & Stepan, A. F. Small-molecule lead-finding trends across the roche and genentech research organizations. J. Med. Chem. 65, 3606–3615 (2022).35138850 10.1021/acs.jmedchem.1c02106
11. Perola E An analysis of the binding efficiencies of drugs and their leads in successful drug discovery programs J. Med. Chem. 2010 53 2986 2997 10.1021/jm100118x 20235539
Perola, E. An analysis of the binding efficiencies of drugs and their leads in successful drug discovery programs. J. Med. Chem. 53, 2986–2997 (2010).20235539 10.1021/jm100118x
12. Lyu J Ultra-large library docking for discovering new chemotypes Nature 2019 566 224 10.1038/s41586-019-0917-9 30728502
Lyu, J. et al. Ultra-large library docking for discovering new chemotypes. Nature 566, 224 (2019).30728502 10.1038/s41586-019-0917-9
13. Sadybekov AA Synthon-based ligand discovery in virtual libraries of over 11 billion compounds Nature 2022 601 452 459 10.1038/s41586-021-04220-9 34912117
Sadybekov, A. A. et al. Synthon-based ligand discovery in virtual libraries of over 11 billion compounds. Nature 601, 452–459 (2022).34912117 10.1038/s41586-021-04220-9
14. Bellmann L Penner P Gastreich M Rarey M Comparison of combinatorial fragment spaces and its application to ultralarge make-on-demand compound catalogs J. Chem. Inf. Model. 2022 62 553 566 10.1021/acs.jcim.1c01378 35050621
Bellmann, L., Penner, P., Gastreich, M. & Rarey, M. Comparison of combinatorial fragment spaces and its application to ultralarge make-on-demand compound catalogs. J. Chem. Inf. Model. 62, 553–566 (2022).35050621 10.1021/acs.jcim.1c01378
15. Neumann A Marrison L Klein R Relevance of the trillion-sized chemical space “explore” as a source for drug discovery ACS Med. Chem. Lett. 2023 14 466 472 10.1021/acsmedchemlett.3c00021 37077402
Neumann, A., Marrison, L. & Klein, R. Relevance of the trillion-sized chemical space “explore” as a source for drug discovery. ACS Med. Chem. Lett. 14, 466–472 (2023).37077402 10.1021/acsmedchemlett.3c00021
16. Sunkari YK Siripuram VK Nguyen T-L Flajolet M High-power screening (HPS) empowered by DNA-encoded libraries Trends Pharmacol. Sci. 2022 43 4 15 10.1016/j.tips.2021.10.008 34782164
Sunkari, Y. K., Siripuram, V. K., Nguyen, T.-L. & Flajolet, M. High-power screening (HPS) empowered by DNA-encoded libraries. Trends Pharmacol. Sci. 43, 4–15 (2022).34782164 10.1016/j.tips.2021.10.008
17. Malo N Hanley JA Cerquozzi S Pelletier J Nadon R Statistical practice in high-throughput screening data analysis Nat. Biotechnol. 2006 24 167 175 10.1038/nbt1186 16465162
Malo, N., Hanley, J. A., Cerquozzi, S., Pelletier, J. & Nadon, R. Statistical practice in high-throughput screening data analysis. Nat. Biotechnol. 24, 167–175 (2006).16465162 10.1038/nbt1186
18. Iversen PW Eastwood BJ Sittampalam GS Cox KL A comparison of assay performance measures in screening assays: Signal window, Z’ factor, and assay variability ratio J. Biomol. Screen. 2006 11 247 252 10.1177/1087057105285610 16490779
Iversen, P. W., Eastwood, B. J., Sittampalam, G. S. & Cox, K. L. A comparison of assay performance measures in screening assays: Signal window, Z’ factor, and assay variability ratio. J. Biomol. Screen. 11, 247–252 (2006).16490779 10.1177/1087057105285610
19. Zhang J-H Chung TDY Oldenburg KR A simple statistical parameter for use in evaluation and validation of high throughput screening assays J. Biomol. Screen. 1999 4 67 73 10.1177/108705719900400206 10838414
Zhang, J.-H., Chung, T. D. Y. & Oldenburg, K. R. A simple statistical parameter for use in evaluation and validation of high throughput screening assays. J. Biomol. Screen. 4, 67–73 (1999).10838414 10.1177/108705719900400206
20. Jadhav A Quantitative analyses of aggregation, autofluorescence, and reactivity artifacts in a screen for inhibitors of a thiol protease J. Med. Chem. 2010 53 37 51 10.1021/jm901070c 19908840
Jadhav, A. et al. Quantitative analyses of aggregation, autofluorescence, and reactivity artifacts in a screen for inhibitors of a thiol protease. J. Med. Chem. 53, 37–51 (2010).19908840 10.1021/jm901070c
21. Fox S High-throughput screening: Update on practices and success J. Biomol. Screen. 2006 11 864 869 10.1177/1087057106292473 16973922
Fox, S. et al. High-throughput screening: Update on practices and success. J. Biomol. Screen. 11, 864–869 (2006).16973922 10.1177/1087057106292473
22. Owen SC Doak AK Wassam P Shoichet MS Shoichet BK Colloidal aggregation affects the efficacy of anticancer drugs in cell culture ACS Chem. Biol. 2012 7 1429 1435 10.1021/cb300189b 22625864
Owen, S. C., Doak, A. K., Wassam, P., Shoichet, M. S. & Shoichet, B. K. Colloidal aggregation affects the efficacy of anticancer drugs in cell culture. ACS Chem. Biol. 7, 1429–1435 (2012).22625864 10.1021/cb300189b
23. Rössler SL Grob NM Buchwald SL Pentelute BL Abiotic peptides as carriers of information for the encoding of small-molecule library synthesis Science 2023 379 939 945 10.1126/science.adf1354 36862767
Rössler, S. L., Grob, N. M., Buchwald, S. L. & Pentelute, B. L. Abiotic peptides as carriers of information for the encoding of small-molecule library synthesis. Science 379, 939–945 (2023).36862767 10.1126/science.adf1354
24. McGovern SL Caselli E Grigorieff N Shoichet BK A Common mechanism underlying promiscuous inhibitors from virtual and high-throughput screening J. Med. Chem. 2002 45 1712 1722 10.1021/jm010533y 11931626
McGovern, S. L., Caselli, E., Grigorieff, N. & Shoichet, B. K. A Common mechanism underlying promiscuous inhibitors from virtual and high-throughput screening. J. Med. Chem. 45, 1712–1722 (2002).11931626 10.1021/jm010533y
25. Feng BY Shelat A Doman TN Guy RK Shoichet BK High-throughput assays for promiscuous inhibitors Nat. Chem. Biol. 2005 1 146 148 10.1038/nchembio718 16408018
Feng, B. Y., Shelat, A., Doman, T. N., Guy, R. K. & Shoichet, B. K. High-throughput assays for promiscuous inhibitors. Nat. Chem. Biol. 1, 146–148 (2005).16408018 10.1038/nchembio718
26. Martin, E. J., Polyakov, V. R., Tian, L. & Perez, R. C. Profile-QSAR 2.0: Kinase virtual screening accuracy comparable to four-concentration IC50s for realistically novel compounds. J. Chem. Inf. Model. 57, 2077–2088 (2017).
27. Keiser MJ Predicting new molecular targets for known drugs Nature 2009 462 175 181 10.1038/nature08506 19881490
Keiser, M. J. et al. Predicting new molecular targets for known drugs. Nature 462, 175–181 (2009).19881490 10.1038/nature08506
28. Svetnik V Random forest: A classification and regression tool for compound classification and QSAR modeling J. Chem. Inf. Comput. Sci. 2003 43 1947 1958 10.1021/ci034160g 14632445
Svetnik, V. et al. Random forest: A classification and regression tool for compound classification and QSAR modeling. J. Chem. Inf. Comput. Sci. 43, 1947–1958 (2003).14632445 10.1021/ci034160g
29. Kitchen DB Decornez H Furr JR Bajorath J Docking and scoring in virtual screening for drug discovery: methods and applications Nat. Rev. Drug Discov. 2004 3 935 949 10.1038/nrd1549 15520816
Kitchen, D. B., Decornez, H., Furr, J. R. & Bajorath, J. Docking and scoring in virtual screening for drug discovery: methods and applications. Nat. Rev. Drug Discov. 3, 935–949 (2004).15520816 10.1038/nrd1549
30. Shoichet BK Virtual screening of chemical libraries Nature 2004 432 862 865 10.1038/nature03197 15602552
Shoichet, B. K. Virtual screening of chemical libraries. Nature 432, 862–865 (2004).15602552 10.1038/nature03197
31. Ma J Sheridan RP Liaw A Dahl GE Svetnik V Deep neural nets as a method for quantitative structure-activity relationships J. Chem. Inf. Model. 2015 55 263 274 10.1021/ci500747n 25635324
Ma, J., Sheridan, R. P., Liaw, A., Dahl, G. E. & Svetnik, V. Deep neural nets as a method for quantitative structure-activity relationships. J. Chem. Inf. Model. 55, 263–274 (2015).25635324 10.1021/ci500747n
32. Sheridan RP Machine Learning and Deep Learning Experimental error, kurtosis, activity cliffs, and methodology: What limits the predictivity of QSAR models? J. Chem. Inf. Model. 2020 10.1021/acs.jcim.9b01067 33022174
Sheridan, R. P. et al. Machine Learning and Deep Learning Experimental error, kurtosis, activity cliffs, and methodology: What limits the predictivity of QSAR models?. J. Chem. Inf. Model.10.1021/acs.jcim.9b01067 (2020).33022174 10.1021/acs.jcim.9b01067
33. Wallach I Heifets A Most ligand-based classification benchmarks reward memorization rather than generalization J. Chem. Inf. Model. 2018 58 916 932 10.1021/acs.jcim.7b00403 29698607
Wallach, I. & Heifets, A. Most ligand-based classification benchmarks reward memorization rather than generalization. J. Chem. Inf. Model. 58, 916–932 (2018).29698607 10.1021/acs.jcim.7b00403
34. Chen L Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening PLOS ONE 2019 14 e0220113 10.1371/journal.pone.0220113 31430292
Chen, L. et al. Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening. PLOS ONE 14, e0220113 (2019).31430292 10.1371/journal.pone.0220113
35. Chuang, K. V. & Keiser, M. J. Comment on “Predicting reaction performance in C–N cross-coupling using machine learning”. Science 362, eaat8603 (2018).
36. Gaieb Z D3R Grand Challenge 3: Blind prediction of protein–ligand poses and affinity rankings J. Comput. Aided Mol. Des. 2019 33 1 18 10.1007/s10822-018-0180-4 30632055
Gaieb, Z. et al. D3R Grand Challenge 3: Blind prediction of protein–ligand poses and affinity rankings. J. Comput. Aided Mol. Des. 33, 1–18 (2019).30632055 10.1007/s10822-018-0180-4
37. Gabel J Desaphy J Rognan D Beware of machine learning-based scoring functions on the danger of developing black boxes J. Chem. Inf. Model. 2014 54 2807 2815 10.1021/ci500406k 25207678
Gabel, J., Desaphy, J. & Rognan, D. Beware of machine learning-based scoring functions on the danger of developing black boxes. J. Chem. Inf. Model. 54, 2807–2815 (2014).25207678 10.1021/ci500406k
38. Cerón-Carrasco JP When virtual screening yields inactive drugs: dealing with false theoretical friends ChemMedChem 2022 17 e202200278 10.1002/cmdc.202200278 35726731
Cerón-Carrasco, J. P. When virtual screening yields inactive drugs: dealing with false theoretical friends. ChemMedChem 17, e202200278 (2022).35726731 10.1002/cmdc.202200278
39. McCloskey K Machine learning on DNA-encoded libraries: A new paradigm for hit-finding J. Med. Chem. 2020 63 8857 8866 10.1021/acs.jmedchem.0c00452 32525674
McCloskey, K. et al. Machine learning on DNA-encoded libraries: A new paradigm for hit-finding. J. Med. Chem. 63, 8857–8866 (2020).32525674 10.1021/acs.jmedchem.0c00452
40. Wenzel J Matter H Schmidt F Predictive multitask deep neural network models for ADME-Tox properties: Learning from large data sets J. Chem. Inf. Model. 2019 59 1253 1268 10.1021/acs.jcim.8b00785 30615828
Wenzel, J., Matter, H. & Schmidt, F. Predictive multitask deep neural network models for ADME-Tox properties: Learning from large data sets. J. Chem. Inf. Model. 59, 1253–1268 (2019).30615828 10.1021/acs.jcim.8b00785
41. Feinberg EN PotentialNet for molecular property prediction ACS Cent. Sci. 2018 4 1520 1530 10.1021/acscentsci.8b00507 30555904
Feinberg, E. N. et al. PotentialNet for molecular property prediction. ACS Cent. Sci. 4, 1520–1530 (2018).30555904 10.1021/acscentsci.8b00507
42. Schindler CEM Large-scale assessment of binding free energy calculations in active drug discovery projects J. Chem. Inf. Model. 2020 60 5457 5474 10.1021/acs.jcim.0c00900 32813975
Schindler, C. E. M. et al. Large-scale assessment of binding free energy calculations in active drug discovery projects. J. Chem. Inf. Model. 60, 5457–5474 (2020).32813975 10.1021/acs.jcim.0c00900
43. Bordogna A Pandini A Bonati L Predicting the accuracy of protein–ligand docking on homology models J. Comput. Chem. 2011 32 81 98 10.1002/jcc.21601 20607693
Bordogna, A., Pandini, A. & Bonati, L. Predicting the accuracy of protein–ligand docking on homology models. J. Comput. Chem. 32, 81–98 (2011).20607693 10.1002/jcc.21601
44. Stokes JM A deep learning approach to antibiotic discovery Cell 2020 180 688 702.e13 10.1016/j.cell.2020.01.021 32084340
Stokes, J. M. et al. A deep learning approach to antibiotic discovery. Cell 180, 688-702.e13 (2020).32084340 10.1016/j.cell.2020.01.021
45. Melo MCR Maasch JRMA de la Fuente-Nunez C Accelerating antibiotic discovery through artificial intelligence Commun. Biol. 2021 4 1 13 10.1038/s42003-021-02586-0 33398033
Melo, M. C. R., Maasch, J. R. M. A. & de la Fuente-Nunez, C. Accelerating antibiotic discovery through artificial intelligence. Commun. Biol. 4, 1–13 (2021).33398033 10.1038/s42003-021-02586-0
46. Skinnider MA A deep generative model enables automated structure elucidation of novel psychoactive substances Nat. Mach. Intell. 2021 3 973 984 10.1038/s42256-021-00407-x
Skinnider, M. A. et al. A deep generative model enables automated structure elucidation of novel psychoactive substances. Nat. Mach. Intell. 3, 973–984 (2021).10.1038/s42256-021-00407-x
47. Muegge I Oloff S Advances in virtual screening Drug Discov. Today Technol. 2006 3 405 411 10.1016/j.ddtec.2006.12.002 38620182
Muegge, I. & Oloff, S. Advances in virtual screening. Drug Discov. Today Technol. 3, 405–411 (2006).38620182 10.1016/j.ddtec.2006.12.002
48. N. Muratov, E. et al. QSAR without borders. Chem. Soc. Rev. 49, 3525–3564 (2020).
49. Zhavoronkov A Deep learning enables rapid identification of potent DDR1 kinase inhibitors Nat. Biotechnol. 2019 37 1038 1040 10.1038/s41587-019-0224-x 31477924
Zhavoronkov, A. et al. Deep learning enables rapid identification of potent DDR1 kinase inhibitors. Nat. Biotechnol. 37, 1038–1040 (2019).31477924 10.1038/s41587-019-0224-x
50. Walters WP Murcko M Assessing the impact of generative AI on medicinal chemistry Nat. Biotechnol. 2020 38 143 145 10.1038/s41587-020-0418-2 32001834
Walters, W. P. & Murcko, M. Assessing the impact of generative AI on medicinal chemistry. Nat. Biotechnol. 38, 143–145 (2020).32001834 10.1038/s41587-020-0418-2
51. Scannell JW Blanckley A Boldon H Warrington B Diagnosing the decline in pharmaceutical R&D efficiency Nat. Rev. Drug Discov. 2012 11 191 10.1038/nrd3681 22378269
Scannell, J. W., Blanckley, A., Boldon, H. & Warrington, B. Diagnosing the decline in pharmaceutical R&D efficiency. Nat. Rev. Drug Discov. 11, 191 (2012).22378269 10.1038/nrd3681
52. Wallach, I., Dzamba, M. & Heifets, A. AtomNet: A Deep Convolutional Neural Network for Bioactivity Prediction in Structure-based Drug Discovery. ArXiv Prepr. ArXiv151002855 1–11 (2015).
53. Gniewek P Worley B Stafford K van den Bedem H Anderson B Learning physics confers pose-sensitivity in structure-based virtual screening. 2021 10.48550/arXiv.2110.15459
Gniewek, P., Worley, B., Stafford, K., van den Bedem, H. & Anderson, B. Learning physics confers pose-sensitivity in structure-based virtual screening.10.48550/arXiv.2110.15459 (2021).10.48550/arXiv.2110.15459
54. Stafford KA Anderson BM Sorenson J van den Bedem H AtomNet PoseRanker: Enriching ligand pose quality for dynamic proteins in virtual high-throughput screens J. Chem. Inf. Model. 2022 62 1178 1189 10.1021/acs.jcim.1c01250 35235748
Stafford, K. A., Anderson, B. M., Sorenson, J. & van den Bedem, H. AtomNet PoseRanker: Enriching ligand pose quality for dynamic proteins in virtual high-throughput screens. J. Chem. Inf. Model. 62, 1178–1189 (2022).35235748 10.1021/acs.jcim.1c01250
55. Hsieh C-H Miro1 marks parkinson’s disease subset and miro1 reducer rescues neuron loss in Parkinson’s models Cell Metab. 2019 30 1131 1140.e7 10.1016/j.cmet.2019.08.023 31564441
Hsieh, C.-H. et al. Miro1 marks parkinson’s disease subset and miro1 reducer rescues neuron loss in Parkinson’s models. Cell Metab. 30, 1131-1140.e7 (2019).31564441 10.1016/j.cmet.2019.08.023
56. Reidenbach AG Multimodal small-molecule screening for human prion protein binders J. Biol. Chem. 2020 295 13516 13531 10.1074/jbc.RA120.014905 32723867
Reidenbach, A. G. et al. Multimodal small-molecule screening for human prion protein binders. J. Biol. Chem. 295, 13516–13531 (2020).32723867 10.1074/jbc.RA120.014905
57. Bon C Discovery of novel trace amine-associated receptor 5 (TAAR5) antagonists using a deep convolutional neural network Int. J. Mol. Sci. 2022 23 3127 10.3390/ijms23063127 35328548
Bon, C. et al. Discovery of novel trace amine-associated receptor 5 (TAAR5) antagonists using a deep convolutional neural network. Int. J. Mol. Sci. 23, 3127 (2022).35328548 10.3390/ijms23063127
58. Stecula A Hussain MS Viola RE Discovery of novel inhibitors of a critical brain enzyme using a homology model and a deep convolutional neural network J. Med. Chem. 2020 63 8867 8875 10.1021/acs.jmedchem.0c00473 32787146
Stecula, A., Hussain, M. S. & Viola, R. E. Discovery of novel inhibitors of a critical brain enzyme using a homology model and a deep convolutional neural network. J. Med. Chem. 63, 8867–8875 (2020).32787146 10.1021/acs.jmedchem.0c00473
59. Su S SPOP and OTUD7A Control EWS–FLI1 protein stability to govern ewing sarcoma growth Adv. Sci. 2021 8 2004846 10.1002/advs.202004846
Su, S. et al. SPOP and OTUD7A Control EWS–FLI1 protein stability to govern ewing sarcoma growth. Adv. Sci. 8, 2004846 (2021).10.1002/advs.202004846
60. Pedicone, C. et al. Discovery of a novel SHIP1 agonist that promotes degradation of lipid-laden phagocytic cargo by microglia. iScience 25, 104170 (2022).
61. Huang C Small molecules block the interaction between porcine reproductive and respiratory syndrome virus and CD163 receptor and the infection of pig cells Virol. J. 2020 17 116 10.1186/s12985-020-01361-7 32727587
Huang, C. et al. Small molecules block the interaction between porcine reproductive and respiratory syndrome virus and CD163 receptor and the infection of pig cells. Virol. J. 17, 116 (2020).32727587 10.1186/s12985-020-01361-7
62. Grygorenko, O. O. et al. Generating multibillion chemical space of readily accessible screening compounds. iScience 23, 101681 (2020).
63. Dandapani S Rosse G Southall N Salvino JM Thomas CJ Selecting, acquiring, and using small molecule libraries for high-throughput screening Curr. Protoc. Chem. Biol. 2012 4 177 191 10.1002/9780470559277.ch110252 26705509
Dandapani, S., Rosse, G., Southall, N., Salvino, J. M. & Thomas, C. J. Selecting, acquiring, and using small molecule libraries for high-throughput screening. Curr. Protoc. Chem. Biol. 4, 177–191 (2012).26705509 10.1002/9780470559277.ch110252
64. Schuffenhauer A Library design for fragment based screening Curr. Top. Med. Chem. 2005 5 751 762 10.2174/1568026054637700 16101415
Schuffenhauer, A. et al. Library design for fragment based screening. Curr. Top. Med. Chem. 5, 751–762 (2005).16101415 10.2174/1568026054637700
65. Jacoby E Key aspects of the novartis compound collection enhancement project for the compilation of a comprehensive Chemogenomics drug discovery screening collection Curr. Top. Med. Chem. 2005 5 397 411 10.2174/1568026053828376 15892682
Jacoby, E. et al. Key aspects of the novartis compound collection enhancement project for the compilation of a comprehensive Chemogenomics drug discovery screening collection. Curr. Top. Med. Chem. 5, 397–411 (2005).15892682 10.2174/1568026053828376
66. Petrova T Chuprina A Parkesh R Pushechnikov A Structural enrichment of HTS compounds from available commercial libraries MedChemComm 2012 3 571 579 10.1039/c2md00302c
Petrova, T., Chuprina, A., Parkesh, R. & Pushechnikov, A. Structural enrichment of HTS compounds from available commercial libraries. MedChemComm 3, 571–579 (2012).10.1039/c2md00302c
67. Macarron R Impact of high-throughput screening in biomedical research Nat. Rev. Drug Discov. 2011 10 188 195 10.1038/nrd3368 21358738
Macarron, R. et al. Impact of high-throughput screening in biomedical research. Nat. Rev. Drug Discov. 10, 188–195 (2011).21358738 10.1038/nrd3368
68. Rogers D Hahn M Extended-connectivity fingerprints J. Chem. Inf. Model. 2010 50 742 754 10.1021/ci100050t 20426451
Rogers, D. & Hahn, M. Extended-connectivity fingerprints. J. Chem. Inf. Model. 50, 742–754 (2010).20426451 10.1021/ci100050t
69. Riniker S Landrum GA Open-source platform to benchmark fingerprints for ligand-based virtual screening J. Cheminformatics 2013 5 26 10.1186/1758-2946-5-26
Riniker, S. & Landrum, G. A. Open-source platform to benchmark fingerprints for ligand-based virtual screening. J. Cheminformatics 5, 26 (2013).10.1186/1758-2946-5-26
70. Ren, F. et al. AlphaFold accelerates artificial intelligence powered drug discovery: Efficient discovery of a novel cyclin-dependent kinase 20 (CDK20) Small Molecule Inhibitor (2022).
71. Assessing structural novelty of the first AI-designed drug candidates to go into human clinical trials. CAS https://www.cas.org/resources/blog/ai-drug-candidates.
72. Kohavi, R. & Wolpert, D. Bias plus variance decomposition for zero-one loss functions. in Proceedings of the Thirteenth International Conference on International Conference on Machine Learning 275–283 (Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 1996).
73. Ferrara P Jacoby E Evaluation of the utility of homology models in high throughput docking J. Mol. Model. 2007 13 897 905 10.1007/s00894-007-0207-6 17487515
Ferrara, P. & Jacoby, E. Evaluation of the utility of homology models in high throughput docking. J. Mol. Model. 13, 897–905 (2007).17487515 10.1007/s00894-007-0207-6
74. Walters WP Namchuk M Designing screens: How to make your hits a hit Nat. Rev. Drug Discov. 2003 2 259 266 10.1038/nrd1063 12669025
Walters, W. P. & Namchuk, M. Designing screens: How to make your hits a hit. Nat. Rev. Drug Discov. 2, 259–266 (2003).12669025 10.1038/nrd1063
75. Inglese J High-throughput screening assays for the identification of chemical probes Nat. Chem. Biol. 2007 3 466 479 10.1038/nchembio.2007.17 17637779
Inglese, J. et al. High-throughput screening assays for the identification of chemical probes. Nat. Chem. Biol. 3, 466–479 (2007).17637779 10.1038/nchembio.2007.17
76. mcule database. https://mcule.com/database/.
77. Screening Collections - Enamine. https://enamine.net/compound-collections/screening-collection.
78. Bruns RF Watson IA Rules for identifying potentially reactive or promiscuous compounds J. Med. Chem. 2012 55 9763 9772 10.1021/jm301008n 23061697
Bruns, R. F. & Watson, I. A. Rules for identifying potentially reactive or promiscuous compounds. J. Med. Chem. 55, 9763–9772 (2012).23061697 10.1021/jm301008n
79. Baell JB Holloway GA New substructure filters for removal of pan assay interference compounds (PAINS) from screening libraries and for their exclusion in bioassays J. Med. Chem. 2010 53 2719 2740 10.1021/jm901137j 20131845
Baell, J. B. & Holloway, G. A. New substructure filters for removal of pan assay interference compounds (PAINS) from screening libraries and for their exclusion in bioassays. J. Med. Chem. 53, 2719–2740 (2010).20131845 10.1021/jm901137j
80. Abagyan R Kufareva I The flexible pocketome engine for structural chemogenomics Methods Mol. Biol. Clifton NJ 2009 575 249 279 10.1007/978-1-60761-274-2_11
Abagyan, R. & Kufareva, I. The flexible pocketome engine for structural chemogenomics. Methods Mol. Biol. Clifton NJ 575, 249–279 (2009).10.1007/978-1-60761-274-2_11
81. Le Guilloux V Schmidtke P Tuffery P Fpocket: An open source platform for ligand pocket detection BMC Bioinformatics 2009 10 168 10.1186/1471-2105-10-168 19486540
Le Guilloux, V., Schmidtke, P. & Tuffery, P. Fpocket: An open source platform for ligand pocket detection. BMC Bioinformatics 10, 168 (2009).19486540 10.1186/1471-2105-10-168
82. Butina D Unsupervised data base clustering based on daylight’s fingerprint and tanimoto similarity: A fast and automated way to cluster small and large data sets J. Chem. Inf. Comput. Sci. 1999 39 747 750 10.1021/ci9803381
Butina, D. Unsupervised data base clustering based on daylight’s fingerprint and tanimoto similarity: A fast and automated way to cluster small and large data sets. J. Chem. Inf. Comput. Sci. 39, 747–750 (1999).10.1021/ci9803381
83. RDKit: Open-Source Cheminformatics.
84. Rarey M Dixon JS Feature trees: A new molecular similarity measure based on tree matching J. Comput. Aided Mol. Des. 1998 12 471 490 10.1023/A:1008068904628 9834908
Rarey, M. & Dixon, J. S. Feature trees: A new molecular similarity measure based on tree matching. J. Comput. Aided Mol. Des. 12, 471–490 (1998).9834908 10.1023/A:1008068904628
85. Stafford K Anderson BM Sorenson J van den Bedem H AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High Throughput Screens. 2021 10.26434/chemrxiv-2021-t6xkj
Stafford, K., Anderson, B. M., Sorenson, J. & van den Bedem, H. AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High Throughput Screens.10.26434/chemrxiv-2021-t6xkj (2021).10.26434/chemrxiv-2021-t6xkj
86. Schroedl S Current methods and challenges for deep learning in drug discovery Drug Discov. Today Technol. 2019 32–33 9 17 10.1016/j.ddtec.2020.07.003 33386100
Schroedl, S. Current methods and challenges for deep learning in drug discovery. Drug Discov. Today Technol. 32–33, 9–17 (2019).33386100 10.1016/j.ddtec.2020.07.003
87. Bender A Mussa HY Glen RC Reiling S Molecular similarity searching using atom environments, information-based feature selection, and a Naïve Bayesian classifier J. Chem. Inf. Comput. Sci. 2004 44 170 178 10.1021/ci034207y 14741025
Bender, A., Mussa, H. Y., Glen, R. C. & Reiling, S. Molecular similarity searching using atom environments, information-based feature selection, and a Naïve Bayesian classifier. J. Chem. Inf. Comput. Sci. 44, 170–178 (2004).14741025 10.1021/ci034207y
88. Trott O Olson AJ AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading J. Comput. Chem. 2010 31 455 461 10.1002/jcc.21334 19499576
Trott, O. & Olson, A. J. AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. J. Comput. Chem. 31, 455–461 (2010).19499576 10.1002/jcc.21334
89. Kingma, D. P. & Ba, J. Adam: A Method for Stochastic Optimization. ArXiv14126980 Cs (2017).
90. Waskom ML seaborn: Statistical data visualization J. Open Source Softw. 2021 6 3021 10.21105/joss.03021
Waskom, M. L. seaborn: Statistical data visualization. J. Open Source Softw. 6, 3021 (2021).10.21105/joss.03021
91. Hunter JD Matplotlib: A 2D graphics environment Comput. Sci. Eng. 2007 9 90 95 10.1109/MCSE.2007.55
Hunter, J. D. Matplotlib: A 2D graphics environment. Comput. Sci. Eng. 9, 90–95 (2007).10.1109/MCSE.2007.55
92. Marineau JJ Discovery of SY-5609: A selective, noncovalent inhibitor of CDK7 J. Med. Chem. 2022 65 1458 1480 10.1021/acs.jmedchem.1c01171 34726887
Marineau, J. J. et al. Discovery of SY-5609: A selective, noncovalent inhibitor of CDK7. J. Med. Chem. 65, 1458–1480 (2022).34726887 10.1021/acs.jmedchem.1c01171
93. Gu, X., BAI, H., Barbeau, O. R. & Besnard, J. Aromatic heterocyclic compound, and pharmaceutical composition and application thereof. (2022).
94. Barbay, J. K., Chakravarty, D., Leonard, K., Shook, B. C. & Wang, A. Phenyl and heteroaryl substituted thieno[2,3-d]Pyrimidines and their use as adenosine A2a receptor antagonists (2010).
95. Bell, A. S., Schreyer, A. M. & Versluys, S. Pyrazolopyrimidine compounds as adenosine receptor antagonists (2019).
96. Soldermann, C. P. et al. Pyrazolo pyrimidine derivatives and their use as MALT1 inhbitors (2019).
97. Feng, S. et al. Tricyclic compounds useful in the treatment of cancer, autoimmune and inflammatory disorders (2023).
98. Heiser, U. & Sommer, R. Inhibitors of glutaminyl cyclase (2020).
99. Cheng, X., Liu, Y., Qin, L., Ren, F. & Wu, J. Beta-lactam derivatives for the treatment of diseases (2023).
100. Wylie, A. A. et al. Therapeutic combinations comprising ubiquitin-specific-processing protease 1 (usp1) inhibitors and poly (adp-ribose) polymerase (parp) inhibitors (2021).
101. Wu, J., Qin, L. & Liu, J. Small molecule inhibitors of ubiquitin specific protease 1 (usp1) and uses thereof 2023).
102. John, S. E. S. & Mesecar, A. D. Broad-spectrum non-covalent coronavirus protease inhibitors (2017).
103. Zavoronkovs, A., Ivanenkov, Y. A. & Zagribelnyy, B. Sars-cov-2 inhibitors having covalent modifications for treating coronavirus infections. (2021).
