
==== Front
ArXiv
ArXiv
arxiv
ArXiv
2331-8422
Cornell University

arXiv:2408.17320v1
2408.17320
1
preprint
Article
BioBricks.ai: A Versioned Data Registry for Life Sciences Data Assets
Gao Yifan
Mughal Zakariyya
Jaramillo-Villegas Jose A.
Corradi Marie
Borrel Alexandre
Lieberman Ben
Sharif Suliman
Shaffer John
Fecho Karamarie
Chatrath Ajay
Maertens Alexandra
Teunis Marc A. T.
Kleinstreuer Nicole
Hartung Thomas
Luechtefeld Thomas
30 8 2024
arXiv:2408.17320v1https://creativecommons.org/licenses/by/4.0/ This work is licensed under a Creative Commons Attribution 4.0 International License, which allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use.
http://arxiv.org/abs/2408.17320v1
nihpp-2408.17320v1.pdf
Researchers in biomedical research, public health, and the life sciences often spend weeks or months discovering, accessing, curating, and integrating data from disparate sources, significantly delaying the onset of actual analysis and innovation. Instead of countless developers creating redundant and inconsistent data pipelines, BioBricks.ai offers a centralized data repository and a suite of developer-friendly tools to simplify access to scientific data. Currently, BioBricks.ai delivers over ninety biological and chemical datasets. It provides a package manager-like system for installing and managing dependencies on data sources. Each 'brick' is a Data Version Control git repository that supports an updateable pipeline for extraction, transformation, and loading data into the BioBricks.ai backend at https://biobricks.ai. Use cases include accelerating data science workflows and facilitating the creation of novel data assets by integrating multiple datasets into unified, harmonized resources. In conclusion, BioBricks.ai offers an opportunity to accelerate access and use of public data through a single open platform.

23 pages, 2 figures
==== Body
pmc
