
==== Front
Plant Commun
Plant Commun
Plant Communications
2590-3462
Elsevier

S2590-3462(24)00195-0
10.1016/j.xplc.2024.100925
100925
Correspondence
BjuIR: A multi-omics database with various tools for accelerating functional genomics research in Brassica juncea
Zhang Linna 1238
Xiao Jinyuan 128
Liang Congyuan 1238
Chen Yifan 124
Yu Changchun 5
Zhao Xinle 12
Li Jiawei 12
Yan Mingli 6
Yang Qian 6
Chen Hao 7
Liu Zhongsong 7
Wan Zhengjie wanzj@mail.hzau.edu.cn
5∗
Yang Zhiquan yang_zq@foxmail.com
124∗∗
Yang Qing-Yong yqy@mail.hzau.edu.cn
1234∗∗∗
1 National Key Laboratory for Germplasm Innovation and Utilization of Horticultural Crops, College of Informatics, Huazhong Agricultural University, Wuhan 430070, China
2 Hubei Key Laboratory of Agricultural Bioinformatics and Hubei Engineering Technology Research Center of Agricultural Big Data, College of Informatics, Huazhong Agricultural University, Wuhan 430070, China
3 National Key Laboratory of Crop Genetic Improvement, Hubei Hongshan Laboratory, Huazhong Agricultural University, Wuhan 430070, China
4 Yazhouwan National Laboratory, Sanya 572025, China
5 National Key Laboratory for Germplasm Innovation and Utilization of Horticultural Crops, College of Horticulture and Forestry Sciences, Huazhong Agricultural University, Wuhan 430070, China
6 Crop Research Institute, Hunan Academy of Agricultural Sciences, Changsha 410125, China
7 College of Agronomy, Hunan Agricultural University, Changsha 410128, China
∗ Corresponding author wanzj@mail.hzau.edu.cn
∗∗ Corresponding author yang_zq@foxmail.com
∗∗∗ Corresponding author yqy@mail.hzau.edu.cn
8 These authors contributed equally to this article

25 4 2024
12 8 2024
25 4 2024
5 8 1009257 1 2024
28 3 2024
16 4 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Published: April 25, 2024
==== Body
pmcDear Editor,

Advances in high-throughput omics technologies, along with methodologies for integrating multi-omics datasets, have substantially enhanced the efficiency of identifying candidate genes in breeding (Gusev et al., 2018; Gupta et al., 2019). However, this process is often complex and laborious. To address this challenge, databases that integrate extensive data and enable convenient and efficient functional genomics studies are being developed (Ma et al., 2021; Yang et al., 2023). Brassica juncea (B. juncea), commonly known as mustard, is an economically significant agricultural species with diverse uses as a vegetable, resilient oilseed crop, and source of distinctively flavored condiments (Yang et al., 2018). This diversity of applications has spurred the accumulation of substantial multi-omics data in fundamental research on mustard, but there has not been a specialized platform to fully harness these data for the genetic improvement of mustard. To address this gap, we have developed BjuIR (Brassica juncea Information Resource, available at https://yanglab.hzau.edu.cn/BjuIR), which integrates the most comprehensive mustard omics datasets to date from over 2000 accessions, including genomic, variomic, transcriptomic, phenomic, and metabolomic data. BjuIR provides sophisticated analyses for these mustard multi-omics datasets with user-friendly interfaces, enabling rapid querying of “variant/gene expression–phenotype” associations for rapid identification of candidate genes and greatly benefiting functional genomics research.

Data and functional modules in BjuIR

BjuIR boasts a rich repository of large-scale multi-omics datasets (Figure 1A), encompassing 40 genome assemblies (Supplemental Table 1), 8 869 856 single nucleotide polymorphisms (SNPs) and short insertions/deletions (InDels) across 1614 accessions (Supplemental Figure 1 and Supplemental Table 2), 941 RNA-seq libraries (Supplemental Table 3), 412 metabolites (Supplemental Table 4), phenotypic data spanning 16 traits from 628 accessions (Supplemental Table 5), and 1841 mustard-centric literature entries. Various analysis methods were applied to fully explore the value of these datasets, and the analysis results are organized and accessible in eight modules within BjuIR (Figure 1B).Figure 1 Overview of BjuIR

(A) Large-scale datasets collected in BjuIR.

(B) Eight modules and their functions in BjuIR.

(C) Comparative genomics analysis between the T84-66.V2.0 and AU213.V1.0 genomes in the “Genomics” module.

(D) eFP viewer displaying tissue-specific expression profiles of the gene BjuVA09G49860 in the “Transcriptomics” module.

(E) “Variation–gene expression–phenotype” associations related to gene BIM1 in the “Multi-omics” module.

(F–J) Identification of novel candidate genes/variants associated with γ-/α-tocopherol content in BjuIR.

(F) GWAS for γ-/α-tocopherol content in seeds. The p-value threshold was set at 6.54e−6 based on 1/n, where n represents the number of independent SNPs (n = 152 884).

(G) Local Manhattan plot of GWAS for γ-/α-tocopherol content and heatmap of linkage disequilibrium (LD) blocks. Dot color represents the degree of LD with the lead SNP.

(H) Genes in the LD block are significantly associated with γ-/α-tocopherol content.

(I) Haplotypes formed by combinations of GWAS SNPs in the coding region of BjuVA06G10820 and its 3-kb upstream flanking region.

(J) Comparison of γ-/α-tocopherol content between accessions with different haplotypes. ∗p < 0.05 (Wilcoxon rank-sum test).

The “Genomics” module provides queries for syntenic relationships between genomes and gene annotation; the “Population” module details accession information and provides selection signals of populations; the “Variations” module allows for the exploration of variants and their associations with phenotypes and gene expression levels; the “Transcriptomics” module features gene expression profiles, co-expression networks, and differential expression analysis; the “Phenomics” module presents phenotype data; the “Metabolomics” module provides metabolite information; the “Multi-omics” module facilitates quick queries for “variation–gene expression–phenotype” associations generated from genome-wide association studies (GWASs), transcriptome-wide association analyses (TWASs), expression quantitative trait loci (eQTL) mapping analyses, colocalization analyses, and summary-based Mendelian randomization; and the “Literature” module supports literature studies on mustard. Each module is equipped with user-friendly interfaces for results visualization, data downloads, and seamless navigation among modules and other databases.

Applications and analysis tools in BjuIR

Comprehensive datasets and thoughtfully designed modules in BjuIR are very useful for those involved in functional genomics and efforts to identify candidate genes/variations.

In the “Genomics” module, users can visualize global genome alignments as a dotplot (Figure 1C) and local genome alignments in Gbrowse in the “Genome synteny” interface (Supplemental Figure 2A) and can query homologous gene clusters by entering a gene ID or gene name in the “Gene cluster” interface (Supplemental Figure 2B). In addition, they can explore annotations of gene families and biological pathways for mustard genes in the “Gene family” and “Pathway” interfaces (Supplemental Figure 2C and 2D).

The “Transcriptomics” module offers gene expression profiles across populations and tissues, co-expression analysis, and comparative transcription analysis, as demonstrated in Supplemental Figure 3. These capabilities are demonstrated using the gene BjuA09.WRI1. By entering the gene BjuVA09G49860 (WRI1) in the “Tissue expression profile” and “eFP” interface, users can access the tissue expression profiles of BjuA09.WRI1 and its homologous genes, presented in a heatmap (Supplemental Figure 3A) and eFP viewer (Figure 1D). Users can also explore gene–gene and lncRNA–mRNA co-expression networks related to BjuA09.WRI1 in the “Co-expression network” interface (Supplemental Figure 3B and 3C) and can access differentially expressed genes/lncRNAs in the “Differential expression” interface (Supplemental Figure 3D and 3E).

The “Population” module provides queries for detailed information on 1614 accessions, including their subpopulations, usages, and origins (Supplemental Figure 4 and Supplemental Table 2). This module enables the querying of selection signals, such as π, Tajima’s D, FST, and cross-population extended haplotype homozygosity (XP-EHH), which are calculated using variations to identify candidate regions and genes that may be under selection (Supplemental Figure 5).

The “Variations” module provides detailed information on genetic variants and assessment of how specific variants or haplotypes affect phenotypes and gene expression (Supplemental Figure 6). For instance, by entering the gene ID “BjuVB08G59610” into the “Variations/Single-locus model” interface, the results page displays annotations for all SNPs and InDels within the gene region (Supplemental Figure 6A–6D). By choosing a particular variant, such as “BB_Chr08:62443776” (Supplemental Figure 6D), users can examine its allele frequency across diverse subpopulations and geographic locations (Supplemental Figure 6E and 6F) and can explore its correlation with phenotypic traits, such as thousand seed weight, and with gene expression levels (Supplemental Figure 6G and 6H).

Integrated analyses of multi-omics data in the “Multi-omics” module vastly improve the efficiency of candidate gene discovery in B. juncea. Here, users can submit a gene name, ID, or trait name to reveal associations between variations, gene expression, and phenotypes. This module includes “variation-–trait” associations identified by GWAS (Supplemental Figure 7A and Supplemental Table 6), “variation–gene expression” associations from eQTL analyses (Supplemental Figure 7B and Supplemental Table 7), and “gene expression–trait” associations identified by TWAS (Supplemental Figure 7C Supplemental Table 8), as well as colocalization analyses (Supplemental Figure 7D–7F and Supplemental Table 9). The reliability of these integrated results is demonstrated by a reproducibility rate of 60.42% in comparison to prior findings (Harper et al., 2020), as shown in Supplemental Table 10. The “variation–gene expression–trait” associations are also viewable in a network format, as demonstrated by entering the “BIM1” gene within the “Multi-omics/Association networks” interface, which showcases all related associations in a visual network (Figure 1E).

The “Literature” module offers advanced search capabilities based on keywords, journal names, and publication years, enabling users to efficiently access research advances related to mustard in a specific field. For instance, by entering the keyword “flowering” in this module, users can retrieve relevant literature on “flowering”, with detailed information displayed in a table. In addition, statistics related to the literature, organized by publication year or journal, are presented visually in line graphs and bar charts. An additional feature is the provision of a keyword co-occurrence network for visualization of trends in studies related to mustard flowering (Supplemental Figure 8).

The “Tools” module incorporates 15 essential bioinformatics analysis tools and applications that support user-initiated analyses, including Gene Ontology enrichment, linkage disequilibrium (LD) calculations, SNP matching for germplasm identification, sequence extraction, primer design, and other functions (Supplemental Figure 9).

Case study: Mining novel candidate variants and genes associated with tocopherol content using BjuIR

We illustrate the utility of BjuIR in mining candidate genes and variations using the example of tocopherol, a crucial vitamin E component vital for seed quality and human nutrition. Tocopherol exists in various forms, such as α-tocopherol, which is recognized for having the highest vitamin E activity in mammals and can be derived from γ-tocopherol (Tucker and Townsend, 2005). When the query “γ-/α-tocopherol content in seed” was initiated in the “Multi-omics/GWAS” interface, the analysis rendered a Manhattan plot that revealed two genomic loci associated with γ-/α-tocopherol content on chromosomes AA_Chr02 and AA_Chr06. The AA_Chr02 locus had been reported previously (Harper et al., 2020), but the locus on AA_Chr06 was a novel finding; it comprised an LD block from 5.88 to 5.95 Mb (Figure 1F and 1G), within which seven genes reside (Figure 1H and Supplemental Table 11). Specifically, BjuVA06G10820, notable for containing the largest number of GWAS SNPs in its coding and 3-kb upstream regions, was identified. Its homolog in Arabidopsis thaliana (AT1G15125) is known to encode S-adenosyl-L-methionine-dependent methyltransferases that participate in the conversion of γ-tocopherol to α-tocopherol (Tavva et al., 2007), suggesting that BjuVA06G10820 could be a prime candidate gene within the AA_Chr06 locus. Haplotype analysis of BjuVA06G10820 through the “Single-locus model” interface in the “Variations” module revealed two prevalent haplotypes (Figure 1I). Accessions carrying Haplotype_1 were significantly associated with reduced γ-/α-tocopherol content compared with Haplotype_2, as shown in Figure 1J. Such findings provide a valuable reference for future breeding strategies aimed at boosting α-tocopherol levels in mustard seeds and underscore the ability of BjuIR to identify candidate genes and variants associated with specific traits.

In conclusion, BjuIR is the most extensive and comprehensive multi-omics database to date for functional genomics research in mustard. Its key features include (1) expedited access to each omics dataset and complete analysis results; (2) rapid mining of candidate genes and variants via robust “variant–gene expression–phenotype” associations; (3) multiple user-friendly, online bioinformatics tools; and (4) navigation-friendly interfaces for efficient data mining. With its rich database and thoughtful design, BjuIR proves to be a highly efficient and convenient platform for functional genomics research and candidate gene identification. Looking ahead, BjuIR will continue to incorporate novel omics data, reinforcing its status as an indispensable platform for furthering functional genomics and genetic improvement in mustard.

Data and code availability

Sources of all datasets are described in the supplemental information. All datasets are available at https://yanglab.hzau.edu.cn/BjuIR/download.

Funding

This research was supported by the 10.13039/501100001809 National Natural Science Foundation of China (32072573 , 31872096 , 32322061 , and 32070559 ); the National Key Research and Development Plan of China (2021YFF1000100, 2023YFD1200102-03 ); the Fundamental Research Funds for the Central University HZAU (2662023XXPY001 ); and the Developing Bioinformatics Platform in Hainan Yazhou Bay Seed Lab (no. JBGS-B21HJ0001 ).

Author contributions

Q.-Y.Y., Z.Y., and Z.W. designed the project. L.Z. and Y.C. collected the datasets. L.Z., C.L., Y.C., J.X., and J.L. performed the bioinformatics analysis. J.X., Y.C., and X.Z. developed the BjuIR database. C.Y. provided the pictures of B. juncea germplasms. L.Z., C.L., Z.Y., and Q.-Y.Y. wrote the manuscript. Q.-Y.Y., Z.Y., Z.W., Z.L., H.C., Q.Y., and M.Y. directed the project. All authors read and approved the manuscript.

Supplemental information

Document S1. Supplemental Figures 1–9

Data S1. Supplemental Tables 1–12

Document S2. Article plus supplemental information

Acknowledgments

We thank the bioinformatics computing platform of the National Key Laboratory of Crop Genetic Improvement, Huazhong Agricultural University, managed by Hao Liu. No conflict of interest is declared.

Published by the Plant Communications Shanghai Editorial Office in association with Cell Press, an imprint of Elsevier Inc., on behalf of CSPB and CEMPS, CAS.

Supplemental information is available at Plant Communications Online.
==== Refs
References

Gupta P.K. Kulwal P.L. Jaiswal V. Association mapping in plants in the post-GWAS genomics era Adv. Genet 104 2019 75 154 31200809
Gusev A. Mancuso N. Won H. Kousi M. Finucane H.K. Reshef Y. Song L. Safi A. Schizophrenia Working Group of the Psychiatric Genomics Consortium Transcriptome-wide association study of schizophrenia and chromatin activity yields mechanistic disease insights Nat. Genet 50 2018 538 548 29632383
Harper A.L. He Z. Langer S. Havlickova L. Wang L. Fellgett A. Gupta V. Kumar Pradhan A. Bancroft I. Validation of an associative transcriptomics platform in the polyploid crop species Brassica juncea by dissection of the genetic architecture of agronomic and quality traits Plant J. 103 2020 1885 1893 32530074
Ma S. Wang M. Wu J. Guo W. Chen Y. Li G. Wang Y. Shi W. Xia G. Fu D. WheatOmics: A platform combining multiple omics data to accelerate functional genomics studies in wheat Mol. Plant 14 2021 1965 1968 34715393
Tavva V.S. Kim Y.-H. Kagan I.A. Dinkins R.D. Kim K.-H. Collins G.B. Increased α-tocopherol content in soybean seed overexpressing the Perilla frutescens γ-tocopherol methyltransferase gene Plant Cell Rep. 26 2007 61 70 16909228
Tucker J.M. Townsend D.M. Alpha-tocopherol: roles in prevention and therapy of human disease Biomed. Pharmacother. 59 2005 380 387 16081238
Yang J. Zhang C. Zhao N. Zhang L. Hu Z. Chen S. Zhang M. Chinese root-type mustard provides phylogenomic insights into the evolution of the multi-use diversified allopolyploid Brassica juncea Mol. Plant 11 2018 512 514 29183772
Yang Z. Wang S. Wei L. Huang Y. Liu D. Jia Y. Luo C. Lin Y. Liang C. Hu Y. BnIR: A multi-omics database with various tools for Brassica napus research and breeding Mol. Plant 16 2023 775 789 36919242
