
==== Front
bioRxiv
BIORXIV
bioRxiv
2692-8205
Cold Spring Harbor Laboratory

39071335
10.1101/2024.07.19.604364
preprint
2
Article
Long-read sequencing transcriptome quantification with lr-kallisto
Loving Rebekah K. http://orcid.org/0000-0001-8725-0376

Sullivan Delaney K. http://orcid.org/0000-0002-8359-6705

Reese Fairlie http://orcid.org/0000-0002-9240-0102

Rebboah Elisabeth http://orcid.org/0000-0003-2273-0189

Sakr Jasmine http://orcid.org/0000-0002-4470-3192

Rezaie Narges http://orcid.org/0000-0001-9326-9954

Liang Heidi Y. http://orcid.org/0000-0002-6190-9723

Filimban Ghassan http://orcid.org/0000-0002-2612-2554

Kawauchi Shimako http://orcid.org/0000-0002-8577-4763

Oakes Conrad http://orcid.org/0000-0002-8936-055X

Trout Diane http://orcid.org/0000-0002-4928-5532

Williams Brian A.
MacGregor Grant http://orcid.org/0000-0001-7598-9501

Wold Barbara J. http://orcid.org/0000-0003-3235-8130

Mortazavi Ali http://orcid.org/0000-0002-4259-6362

Pachter Lior http://orcid.org/0000-0002-9164-6231

09 9 2024
2024.07.19.604364https://creativecommons.org/licenses/by/4.0/ This work is licensed under a Creative Commons Attribution 4.0 International License, which allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use.
http://biorxiv.org/lookup/doi/10.1101/2024.07.19.604364
nihpp-2024.07.19.604364.pdf
RNA abundance quantification has become routine and affordable thanks to high-throughput “short-read” technologies that provide accurate molecule counts at the gene level. Similarly accurate and affordable quantification of definitive fulllength, transcript isoforms has remained a stubborn challenge, despite its obvious biological significance across a wide range of problems. “Long-read” sequencing platforms now produce data-types that can, in principle, drive routine definitive isoform quantification. However some particulars of contemporary long-read datatypes, together with isoform complexity and genetic variation, present bioinformatic challenges. We show here, using ONT data, that fast and accurate quantification of long-read data is possible and that it is improved by exome capture. To perform quantifications we developed lr-kallisto, which adapts the kallisto bulk and single-cell RNA-seq quantification methods for long-read technologies.
==== Body
pmc
