
==== Front
G3 (Bethesda)
Genetics
g3journal
G3: Genes | Genomes | Genetics
2160-1836
Oxford University Press US

39099140
10.1093/g3journal/jkae161
jkae161
Investigation
AcademicSubjects/SCI01180
AcademicSubjects/SCI01140
Genome-wide association studies from spoken phenotypic descriptions: a proof of concept from maize field studies
https://orcid.org/0000-0002-9202-1911
Yanarella Colleen F Department of Agronomy, Iowa State University, Ames, IA 50011, USA
Bioinformatics and Computational Biology Program, Iowa State University, Ames, IA 50011, USA

https://orcid.org/0000-0001-9806-7534
Fattel Leila Department of Agronomy, Iowa State University, Ames, IA 50011, USA
Interdepartmental Genetics and Genomics Program, Iowa State University, Ames, IA 50011, USA

https://orcid.org/0000-0003-0069-1430
Lawrence-Dill Carolyn J Department of Agronomy, Iowa State University, Ames, IA 50011, USA
Bioinformatics and Computational Biology Program, Iowa State University, Ames, IA 50011, USA
Interdepartmental Genetics and Genomics Program, Iowa State University, Ames, IA 50011, USA
Department of Genetics, Development and Cell Biology, Iowa State University, Ames, IA 50011, USA
College of Agriculture and Life Sciences, Iowa State University, Ames, IA 50011, USA

Hufford M Editor
Corresponding author: Department of Soil and Crop Sciences, Colorado State University, Fort Collins, CO 80523, USA. Email: carolyn.lawrence-dill@colostate.edu Present address: Department of Soil and Crop Sciences, Colorado State University, Fort Collins, CO 80523, USA
Conflicts of interest The author(s) declare no conflicts of interest.

9 2024
05 8 2024
05 8 2024
14 9 jkae16112 5 2024
23 6 2024
05 8 2024
© The Author(s) 2024. Published by Oxford University Press on behalf of The Genetics Society of America.
2024
https://creativecommons.org/licenses/by/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.

Abstract

We present a novel approach to genome-wide association studies (GWAS) by leveraging unstructured, spoken phenotypic descriptions to identify genomic regions associated with maize traits. Utilizing the Wisconsin Diversity panel, we collected spoken descriptions of Zea mays ssp. mays traits, converting these qualitative observations into quantitative data amenable to GWAS analysis. First, we determined that visually striking phenotypes could be detected from unstructured spoken phenotypic descriptions. Next, we developed two methods to process the same descriptions to derive the trait plant height, a well-characterized phenotypic feature in maize: (1) a semantic similarity metric that assigns a score based on the resemblance of each observation to the concept of ‘tallness’ and (2) a manual scoring system that categorizes and assigns values to phrases related to plant height. Our analysis successfully corroborated known genomic associations and uncovered novel candidate genes potentially linked to plant height. Some of these genes are associated with gene ontology terms that suggest a plausible involvement in determining plant stature. This proof-of-concept demonstrates the viability of spoken phenotypic descriptions in GWAS and introduces a scalable framework for incorporating unstructured language data into genetic association studies. This methodology has the potential not only to enrich the phenotypic data used in GWAS and to enhance the discovery of genetic elements linked to complex traits but also to expand the repertoire of phenotype data collection methods available for use in the field environment.

spoken descriptions
association studies
phenotyping
maize
genes
plant height
Iowa State University Plant Sciences Institute DGE 10.13039/100000082 1545453 NSF 10.13039/100000001 #2021-67021-35329
==== Body
pmcIntroduction

Collecting phenotype data can be slow, which limits the speed of association genetics and genomic studies for trait improvement. High-throughput phenotyping methods are an area of development that concentrates on engineering sensors and unmanned vehicles to collect (mainly visual) data about traits of various crop species (reviewed in Yang et al. 2020). These methods are beneficial for collecting large amounts of data in an automated fashion, but there are difficulties in deploying these tools in a field environment, and some traits are not detectable by images alone. Additionally, manually collecting phenotypes with pen-and-paper or tablet-and-stylus is time-consuming and generally requires predefined traits of interest. Sensors, imaging, and barcodes make data organization easier for large quantities of data (Kazic 2020; Yao et al. 2021; Sarić et al. 2022). An underdeveloped area of in-field phenotyping ripe for exploration is using natural language descriptions of plants. Platforms exist where audio descriptions are recorded (Kazic 2020). However, the biologically relevant data in spoken phenotypes thus far has remained inaccessible for association studies and other applications.

As demonstrated by Oellrich et al. (2015), language-based descriptions of phenotypes can be analyzed computationally to recover known biological associations in plants. Their work used phenotype descriptions derived from various ontologies, then structured the data as entity quality (EQ) statements (where an entity is a feature, e.g. whole plant, and quality is a describer, e.g. dwarf-like), which results in less intensive computations (Mungall et al. 2010). Semantic or word-meaning similarity methods have also shown promise in ascertaining biologically meaningful genetic associations without the burden of manual curation (Braun and Lawrence-Dill 2020; Braun et al. 2020). Additionally, pretrained models have enabled free-text descriptions of plant phenotypes for association studies based on semantics (Braun et al. 2021). The developments in the computational processing of natural language plant phenotypes through unstructured text formats contribute to conceptualizing methods for recording spoken descriptions of phenotypes.

Given the various difficulties associated with collecting structured phenotype data for trait analyses alongside recent advances in semantic reasoning by computer systems, we were curious to find out whether computing on unstructured descriptions of phenotype might now be tractable for biological inferences in a field setting. Imagine collecting phenotype data simply by walking through a field and using spoken words and vocabularies to describe phenotypes, then taking those audio recordings back to the lab as input along with genomics data to run GWAS. Are the analytical tools advanced enough for this kind of approach to phenotypic data collection and analysis? This question is the basis for experiments and results presented here.

To test whether associations between genomic variants and phenotypes could be derived from unstructured spoken language, we reasoned that a well-characterized diversity panel that had many associations already derived would be required as ground truth, which led us to select the Wisconsin Diversity (WiDiv) panel for this work. WiDiv (1) was developed to grow in the upper midwestern region of the United States; (2) includes only lines with similar flowering time and phenology, and (3) is made up of lines that have both genetic and phenotypic diversity (Hansey et al. 2011). Researchers using the WiDiv panel have increased the included accessions, investigated flowering time and biomass yield traits, and generated genetic marker data for the panel (Hansey et al. 2011; Hirsch et al. 2014; Mazaheri et al. 2019; Mural et al. 2022b).

Because the WiDiv panel data has expansive genotypic and phenotypic trait data available, we collected spoken descriptions of phenotypes for numerous traits (height, leaf width, stalk, tassel, leaf, braceroot, and overall color, etc.) during the summer of 2021 (Yanarella et al. 2023a). The objectives of this research were to (1) detect phenotype descriptions from spoken language, (2) demonstrate techniques for extracting phenotype data for a proof of concept analysis involving the plant height trait for genome-wide association studies (GWAS), (3) perform GWAS with phenotypes derived from speech, and (4) use available gene function data to review and assess known and novel gene trait associations.

The methodologies outlined in this article are versatile enough to analyze a range of traits present in the datasets; however, our demonstration centered on plant height due to its well-established genetic associations. It is important to acknowledge the subjective element introduced by using spoken, natural language descriptions, which diverges from the strict reproducibility and comparability offered by measured phenotypes and controlled vocabularies. Despite this, the ease of data gathering could prove beneficial, especially when it leads to the discovery of previously unknown gene-trait associations. Our successful identification of known genetic associations to the plant height trait validates this novel approach and sets a precedent for the future of phenotypic data acquisition and analysis. This study envisions a future where one can simply walk into a field, articulate observations, and receive trait-gene associations in return, leveraging the natural language as a powerful analytical tool.

Materials and methods

We used a genotypic dataset that includes WiDiv panel taxa (lines). A dataset of 18 million SNP markers (Mural et al. 2022a) obtained from aggregating RNA-Seq and resequencing techniques for 1,051 taxa (described in Mural et al. 2022b). Phenotypic datasets (described in Yanarella et al. 2024), include audio recordings as well as manual measurements of plant height for 686 unique WiDiv panel taxa (Yanarella et al. 2023a), hereafter referred to as the Yanarella et al. dataset. This dataset contains an additional 25 taxa (Supplementary Table 1) that were positive controls for the analysis of spoken descriptions of phenotypes, as these genetics lines were expected to have noticeable and readily describable phenotypes. These positive control taxa are not members of the WiDiv panel. Informed consent for the spoken data from participants was collected per Iowa State University’s Institutional Review Board’s (IRB) Exempt Project status, and volunteers provided informed consent for using their spoken observations. Phenotypic data include measurements for plant height and spoken descriptions of plants grown in a field environment. The Mural et al. and Yanarella et al. datasets share 653 taxa (Fig. 1), which were used for further analyses described here.

Fig. 1. Comparison of intersections of Mural et al. and Yanarella et al. WiDiv dataset taxa (positive controls not included), where n is the number of unique taxa in each dataset. Mural et al. dataset (in blue, n=1,051), and Yanarella et al. dataset (in orange, n=686).

Venn Diagram. Mural dataset n=1,051. Yanarella dataset n=686. Overlap is 653.

Spoken phenotype collection summary

We reported on phenotype descriptions recorded by deidentified student workers as a peer-reviewed dataset elsewhere (Yanarella et al. 2024). We refer to the Yanarella et al. dataset throughout, and detailed information on data collection and dataset formats can be accessed via Yanarella et al. (2024). These data and the methods used for their collection are reviewed in brief here for context.

Unstructured spoken descriptions were exclusively captured in the field environment. The field in which the recordings were taken included two replicates planted in a randomized incomplete block design. The first block consisted of 31 WiDiv panel taxa for a seed increase, and the second block consisted of 8 B73 experimental control rows, 25 positive control taxa, and 655 unique WiDiv panel taxa. The second block contained two rows of the WiDiv panel line MEF156-55-2. Therefore, the recordings were taken over 720 rows in each replicate. Manually collected phenotypes included plant height as well as ear and tassel features. Phenotype data were collected during July and August of 2021, a time when plants had reached maturity.

Each of the deidentified student workers selected NATO code names. The students who were undergraduate Agronomy, Biology, and Genetics Majors at Iowa State University are known as “Delta”, “Golf”, “Kilo”, “Lima”, “Mike”, “Quebec”, “Victor”, “Yankee”, and “Zulu”. Each participant was instructed to state their NATO code name and the row tag number before observing the plants in each row (Fig. 2a). This procedure ensured the participant’s deidentified connection to the row number and spoken observation while enabling the parsing of each observation so that multiple row observations could be recorded in the same file (data available from Yanarella et al. 2023a, and described in Yanarella et al. 2024). Participants were asked to comment on overall appearance, including color and height, tassel, ear, leaves, braceroots, and any disease. Two examples of how a given row might be described are as follows. “Alpha, 1,641. Plants are flowering, there’s a lot of variability in height, tassels have emerged, the anthers have everted, brace roots at two nodes, some braceroots point upward”. Another: “Alpha, 2,287. Plants are tall. They have brace roots at up to three nodes. Those brace roots are chunky. They have red to purple anthocyanins in rings. Leaves are very much upright. Silks are yellow, tassels are yellow”. Because many individuals passed each row many times, each row has a diversity of descriptions in the full dataset.

Fig. 2. Spoken phenotype process overview. a) In field spoken phenotype descriptions collection, b) spoken phenotype data processing multiple ways, including transcript production and methods for generating numeric representations of phenotypes for traits, and c) top box: detecting phenotypes from positive control accession descriptions. Lower two boxes: GWAS using data derived from spoken observations.

Phenotype detection from spoken descriptions of positive control lines

A subset of spoken observation transcripts collected in the field containing the positive control accessions were parsed (Fig. 2b). Four to six terms from the description and phenotype records for each accession were drawn from MaizeGDB (Woodhouse et al. 2021). One term from these lists was used to collect synonyms from Merriam-Webster (Merriam-Webster 2023) and WordHippo (Kat IP Pty Ltd 2008) thesaurus services (Table 1 and Fig. 2c-top). The number of rows containing at least one synonym related to each accession’s descriptions and phenotype records was calculated as a proportion to the number of observations for that accession.

Table 1. Proportion of word usage for describing positive control accessions.

Accession	Rows observed count	Merriam-Webster	WordHippo	
		Number of synonyms	Proportion	Number of synonyms	Proportion	
528A Hsf1-N1595	53	32	0.2264	85	0.7358	
528AA Hsf1-N2559	55	32	0.2000	85	0.6727	
528J Hsf1-N1603	54	32	0.2778	85	0.8148	
X28F cr4-6143	54	34	0.4815	83	0.5000	
X29C cr4-N590C	54	34	0.5370	83	0.4630	
X29D cr4-N647	54	34	0.6296	83	0.5741	
702B o2 v5 ra1-Ref gl1	53	34	0.3208	73	0.4528	
703D o2 ra1-Ref gl1	54	34	0.4074	73	0.5185	
711H ra1-N408E	54	34	0.5000	73	0.5370	
117D tb1	53	32	0.2075	71	0.3208	
117DA tb1-8963	54	32	0.3889	71	0.4630	
X03A sr3	55	32	0.5455	22	0.4364	
X03J sr3-N504A	54	32	0.5926	22	0.6296	
104B rli1-N2302A	54	49	0.1667	77	0.1667	
104C rli1-N2276	55	49	0.2182	77	0.2000	
703J Rs1-O	54	49	0.1667	77	0.2037	
703K Rs1-Z	54	49	0.1667	77	0.2593	
219L B1-S; R1-r pl1-McClintock	54	28	1.0000	28	1.0000	
617A pl1-0 c1-EMS Sh1 B1-S R1-r	55	28	0.2182	28	0.2727	
M241C A1 A2 B1 C1 C2 Pl1 Pr1 R1-r a	54	28	1.0000	28	1.0000	
M341B A1 A2 B1 C1 C2 pl1 Pr1 R1-r	55	28	0.8364	28	0.8364	
U740G Fbr1-N1602	54	31	0.0000	6	0.0370	
916G Trn1-N1597	54	28	0.9815	57	0.9815	
708C o15-N1117	54	71	0.3333	120	0.7037	
119E Ts6	54	43	0.2778	47	0.8889	
a Accession used for example illustrated in Fig. 4.

Preprocessing phenotypic datasets

Plant height measurement data

R Scripts (v.4.2.2 and v.4.3.1) (R Core Team 2023) were developed to process the measuring and scoring data from Yanarella et al. dataset such that only plant height observations were retained for each of the three observation sets (each row in the field was measured and scored by two teams made up of participants and one observation set was collected by a volunteer). Replicate number was programmatically added to these data, and positive controls were removed. The 653 taxa shared between the datasets were retained.

Best Linear Unbiased Estimators (BLUE) values were calculated for each taxon using R’s built-in lm function to perform linear regressions and the emmeans v.1.8.7 package (Lenth 2023), where taxa and replicate were fit as fixed effects. Best Linear Unbiased Prediction (BLUP) corrected mean values were calculated for each taxon with the lmer function of the lme4 v.1.1-34 package (Bates et al. 2015) where taxa, replicate, and row number were fit as random effects. Visualization of diagnostics plots of the models (Supplementary Figs. 1, 2) were generated using the ggResidpanel v.0.3.0 package (Goode and Rey 2022).

Semantic similarity for plant height spoken data

Text transcripts of spoken data from the Yanarella et al. dataset were processed using Python v.3.8.2 (Van Rossum and Drake 2009). The spaCy v.3.5.1 package (Honnibal and Montani 2023) and the TensorFlow v.2.12.0 (Abadi et al. 2016) spaCy Universal Sentence Encoder v.0.4.6 (Mensio 2023) were used to process the transcripts to obtain semantic similarity scores. Three phrases, “tall”, “tall plant”, and “tall height”, were compared to each row observation through spaCy’s similarity function using the pretrained large English universal sentence encoder (en_use_lg) from TensorFlow. A dataset of similarity scores in the form of values from 0 to 1 was generated (Fig. 2b). The spaCy Universal sentence encoder and the similarity function we used scores semantic similarity from 0 to 1 based on the phrase used to compare the observations. Therefore, we selected to use all tall terms as the query (rather than also focusing on phrases for short stature) so that observations for tall plants would be closer to a score of 1 (observations meaning the opposite, i.e. small plants would be scored closer to 0). If the word short were substituted for tall as the query, then smaller would be scored closer to 1 and tall plants closer to 0.

Similarity scores for the 653 taxa shared by both datasets were retained, encompassing 35,709 rows or 91.92% of the original rows observed (Table 2). The similarity scores for the “tall” query were used as input to calculate BLUEs and BLUPs in the same manner as described in the Plant Height Measurement Data section, and visualizations of diagnostics plots of the models (Supplementary Figs. 3, 4) were generated.

Table 2. Observation retention by spoken phenotype method.

Participant	Total rows observed counta	Spoken phenotype processing method	
		Tall queryb	Manual binsc	
Delta	4,326	3,974 (0.9186)	3,911 (0.9041; 0.9841)	
Golf	4,317	3,969 (0.9194)	3,711 (0.8596; 0.9350)	
Kilo	4,321	3,973 (0.9195)	3,960 (0.9165; 0.9967)	
Lima	4,322	3,975 (0.9197)	3,747 (0.8670; 0.9426)	
Mike	4,302	3,953 (0.9189)	3,937 (0.9152; 0.9960)	
Quebec	4,308	3,959 (0.9190)	3,241 (0.7523; 0.8186)	
Victor	4,329	3,980 (0.9194)	3,946 (0.9115; 0.9915)	
Yankee	4,308	3,960 (0.9192)	3,943 (0.9153; 0.9957)	
Zulu	4,314	3,966 (0.9193)	3,813 (0.8839; 0.9614)	
Sum	38,847	35,709 (0.9192)	34,209 (0.8806; 0.9580)	
a Rows observed count includes data collected for all taxa within the Yanarella et al. dataset, including 25 positive control taxa and 33 taxa not found in the Mural et al. dataset.

b The count of row observations utilized for semantic similarity for spoken data methods represents the data 653 intersecting taxa between Yanarella et al. and Mural et al., and the data in parentheses are the proportion of data retained from total rows observed.

c The count of row observations utilized for binning for spoken data methods represents the data 653 intersecting taxa between Yanarella et al. and Mural et al., and the data in parentheses are the proportion of data retained from total rows observed (left) and the proportion of data retained from the 653 intersecting taxa which have plant height terms (right).

Binning for plant height spoken data

The transcripts of spoken plant descriptions were mined manually and programmatically for phrases directly related to narrations about plant height. A set of 797 plant height phrases were identified and manually curated and binned from 0 to 7, where 0: no growth, 1: very short plants, 2: short plants, 3: short-medium height plants, 4: medium height plants, 5: medium-tall height plants, 6: tall plants, 7: very tall plants. Bin values were assigned to observations for the 653 taxa shared by both datasets (Fig. 2b). Of the total text transcripts, there were 34,209 or 88.06% row observations that were retained and binned (Table 2).

BLUEs and BLUPs were calculated as described in the Plant Height Measurement Data section, and visualizations of diagnostics plots of the models (Supplementary Figs. 5, 6) were generated. Additionally, the R nnet v.7.3-19 package (Venables and Ripley 2002) was used to perform multinomial logistic regression to predict height phrase bins, where taxa and replicate were fit as fixed effects.

Preprocessing the genotypic dataset

To ease the subsetting step to obtain the data for the 653 that overlap between the Mural and Yanarella datasets, Trait Analysis by aSSociation, Evolution and Linkage (TASSEL) Version 5.0 Standalone (Bradbury et al. 2007) was used to convert the Mural et al. genotypic data from a variant call format (vcf) formatted file to a HapMap (hmp) formatted file. The data were then processed to contain marker information for the 653 taxa shared with the Yanarella et al. dataset, these data were grouped by chromosome (because datasets were so large as to exceed the limits of the HPC memory we were able to utilize), and HapMap files were generated for each chromosome. The chromosome files were sorted by maker position from lowest position to highest position.

The sorted chromosome files were reformatted to vcf files using TASSEL Version 5.0 Standalone, then vcftools v.0.1.14 (Danecek et al. 2011) concatenated these files, and the resulting file was zipped. The concatenated data (a subset of 653 lines) were then again converted to a VCF format for use with PopLDdecay v3.42 (Zhang et al. 2018) to analyze and visualize Linkage Disequilibrium (LD) Decay of the Mural et al. genotypic data.

GWAS and analysis

Genome Association and Prediction Integrated Tool (GAPIT) 3 v.3.1, 2022.4.16 (Lipka et al. 2012; Tang et al. 2016; Wang and Zhang 2021) was used to perform Fixed and random model Circulating Probability Unification (FarmCPU) (Liu et al. 2016) and Mixed Linear Model (MLM) (Yu et al. 2005) on each of the phenotypic datasets using the Mural et al. marker dataset for genotypic input. Each chromosome was run individually, and PCA.total parameter was set to 3 for all analyses (Fig. 2c). This manuscript focuses on the FarmCPU analyses, the MLM processing and results are available as described in Data availability section.

Manhattan plots from the resulting GAPIT analyses were generated using the ggplot2 v.3.4.3 package (Wickham 2016). We used the RAINBOWR v.0.1.29 package’s (Hamazaki and Iwata 2020) CalcThreshold to determine the Bonferroni threshold for each analysis with a sig.level of 0.05. SNPs that were identified as above the Bonferroni threshold for each analysis were viewed on MaizeGDB’s implementation of GBrowse2 (Generic Genome Browser v.2.55) (Stein 2013) using Maize B73 RefGen_v4; gene IDs were collected within ±300 kilobases (kb), based on the LD decay curve generated by PopLDdecay, of the identified SNPs (Supplementary Fig. 7).

We collected a list of genes shown to influence plant height (Table 3, Supplementary Table 2) and compared the gene IDs from the GWAS analyses using a web-based intersection and Venn diagram tool (Sterck 2021) to determine if these previously published plant height genes were identified within the ±300 kb region indicated by the LD decay curve for each of our analyses. Additionally, gene ontology (GO) terms for the gene IDs within ±300 kb of SNPs identified as significant were obtained using the B73 RefGen_V4 Zm00001d.2 annotations generated by Maize GO Annotation—Methods, Evaluation, and Review (maize-GAMER) tool (Wimalanathan and Lawrence-Dill 2017; Wimalanathan et al. 2018) and the R package GO.db v.3.17.0 (Carlson 2023) was used to collect terms associated to the GO IDs (Supplementary Table 3). Where no GO terms were associated with genes or gene models or those associated were not clearly related to the plant height trait, knowledge of gene function via associated Arabidopsis thaliana orthologs listed in MaizeGDB and linked to TAIR were reviewed for possible insight (Andorf et al. 2016; Reiser et al. 2022).

Table 3. Plant height gene models count identified from publications.

Publication	Plant height related gene model counta	
Jansson (1994)	1	
Winkler and Helentjaris (1995)	1	
Austin and Lee (1996)	3	
Multani et al. (2003)	2	
Blakeslee et al. (2005)	1	
Geisler and Murphy (2005)	1	
Brooks et al. (2009)	1	
Wu et al. (2009)	7	
Bai et al. (2009)	6	
Lawit et al. (2010)	2	
Hartwig et al. (2011)	1	
Salvi et al. (2011)	4	
Weng et al. (2011)	1	
Teng et al. (2012)	1	
Gallavotti (2013)	1	
Peiffer et al. (2014)	7	
Wallace et al. (2016)	312	
Mazaheri et al. (2019)	10	
Azodi et al. (2020)	876	
Mural et al. (2022b)	347	
a Count of gene model IDs identified for Zm00001d.2 gene model and Zm-B73-REFERENCE-GRAMENE-4.0 assembly version.

Results and discussion

Student participants made three complete passes of the field (∼4,320 observations per student participant (Table 2, total rows observed count) and used their individual wording and phraseology to describe phenotypes. Recorded observations varied from 2 to 241 words in length (Fig. 3). From this dataset (described fully in Yanarella et al. 2024), all analyses and results reported here are derived.

Fig. 3. Distributions of each student participant’s word count per observation. X axis shows number of words for individual row observation. Y axis indicates each participant’s NATO name. Boxplots are included within each of the violin plots, individual outlier points not represented.

Visually striking phenotypes are detected from unstructured spoken descriptions

We reasoned generally that including visually observable phenotypes with known phenotype–gene associations would be a first step toward understanding whether individuals without specific training on describing plant phenotypes could recover known phenotype–gene associations. To explore student participants’ ability to identify and describe phenotypes, 25 positive control accessions that, if grown in the appropriate environmental conditions, would show visually “dramatic” phenotypes were included in the field.

The 25 positive control accessions were observed 53–55 times over all nine student participants (Table 1). We utilized Merriam-Webster (2023) and WordHippo (Kat IP Pty Ltd 2008) thesaurus services to determine participants’ ability to identify words synonymous with descriptors that characterize positive control phenotypes as demonstrated in Fig. 4a–c. For example, accessions M241C A1 A2 B1 C1 C2 Pl1 Pr1 R1-r and 219L B1-S; R1-r pl1-McClintock (gene name colored1 and colored plant1, Supplementary Table 1) had at least one synonym in each of the observations made by the student participants as indicated in Table 1 by the proportion of 1.000 for both Merriam-Webster and WordHippo synonyms. While accessions U740G Fbr1-N1602 (gene name few branched1), 703J Rs1-O 1 and 703K Rs1-Z (gene name rough sheath1, Supplementary Table 1) had low proportions of observations having at least one synonym for each observation as indicated in Table 1.

Fig. 4. Example of detecting phenotypes from positive control accession descriptions. a) Transcript of a spoken description for positive control accession M241C A1 A2 B1 C1 C2 PI1 Pr1 R1-r, b) image of positive control accession M241C A1 A2 B1 C1 C2 PI1 Pr1 R1-r, c) example of phenotype description terms, and synonyms from Merriam-Webster and WordHippo, to demonstrate a description having at least one instance of a synonymous word for the phenotype of interest.

Our findings indicate that nonexpert participants, unaware of expected phenotypes, can identify and describe what they observe using terms that could be valuable for genomic association studies. These findings are, however, constrained by the number of synonyms used for phenotype descriptions and the specific environmental conditions of the plant growth. The reliance on thesaurus services could also potentially reduce the proportion of observations with at least one synonym, particularly if participants opted for informal descriptions. Furthermore, accurate descriptions of our intended positive control phenotypes were contingent upon favorable field environment and weather conditions.

Plant height assessments can be derived from unstructured spoken descriptions

Given that visually apparent phenotypes were described successfully by nonexpert participants for control lines grown in the field, we moved on to assessing whether plant height, a continuous and subjective trait from “short” to “tall”, could be derived from the WiDiv population, a large association mapping panel.

Two methods were employed to preprocess text transcriptions of spoken descriptions of plants. The semantic similarity method of comparing the term “tall” to each row observation retained 91.92% of the full set of row recordings captured (including the 25 positive controls and 33 accessions unique to the Yanarella et al. dataset) by the student participants, demonstrating that 35,709 row observations were made with taxa in both datasets (Table 2). The manual bin method of identifying phrases related to plant height and binning them based on apparent semantic similarity retained 86.06% of the full set of row recordings captured by the student participants and 95.80% of the observations with plant height phrases made with taxa in both datasets, which results in 34,209 row observations for manually binned data (Table 2).

Both methods parse information about the plant height trait and process the data into a format appropriate as input into available GWAS tools and models. Using a query term for plant height and semantic similarity requires less manual curation and was implemented on a larger subset of data. The benefit of the binning method is that it reduces the noise (where noise in the context of natural language processing refers to additional information not directly related to the particular topic of interest). Only observations with plant height-related terms were considered.

A limitation of the query term and semantic similarity method is retaining noisy data because this method compares the “tall” query to each observation and relies on pretrained models. An example of noise comes from the participant NATO code name “Victor”. Their recording for row 1,456 on July 16, 2021, tall and height green all the way to the bottom … super short hairs on top that are quite prickly … in general there are the brace roots are short fat and light green in color, in which the similarity function of spaCy University Sentence Encoder (Mensio 2023) when compared to the “tall” query string determine the semantic similarity score as 0.0848. The shortcomings of the binning method are the time-consuming nature of curating lists of phrases relevant to the trait of interest and the loss of data where phrases directly related to plant height were not specified.

Genomic associations for plant height can be derived from unstructured spoken descriptions of phenotype

Efforts to assess whether spoken data could be used for GWAS were structured as follows. We collected a dataset of manually measured plant heights, which would serve as ground truth for assessing any language-based findings. We also processed spoken descriptions of lines in two ways (semantic similarity versus a binning method), and carried out GWAS using BLUEs to model the parameters taxa and replicate as fixed effects and BLUPs to model the parameters taxa, replicate, and row number as random effects. Once regions of the genome associated with the trait plant height were derived for manually measured, semantic similarity, and binned datasets, significant regions that included genes known to be involved in plant height were highlighted. Loci identified as significant from both spoken datasets were compared to the manual, ground truth dataset to determine whether the same regions were identified across both data collection methods, and all genes in regions of significant association for the plant height trait were assessed for GO terms associated with plant growth hormones auxin, brassinosteroid, and gibberellin.

Ground truth: manual plant height measurement GWAS

We performed association studies using FarmCPU (Liu et al. 2016) on three categories of phenotype data. The first phenotype category was ground truth (measured) plant height data using BLUEs (Fig. 5a) and BLUPs (Fig. 5b). For the BLUE analysis, 21 significant SNPs above the Bonferroni threshold of 8.55 were identified, and of those SNPs, we discovered 10 (Supplementary Table 3; colored orange) in which at least one plant height gene was detected in the literature within ±300 kb (Supplementary Table 2). The BLUP analysis identified 29 significant SNPs, 9 (Supplementary Table 3) where at least one plant height gene was discovered in the literature within ±300 kb (Table 3, Supplementary Table 2).

Fig. 5. Manually measured height—phenotypic data. Manhattan plot generated using GAPIT and FarmCPU using measured height data a) BLUEs and b) BLUPs using Mural et al. genotypic data. The horizontal red dashed line indicates the Bonferroni threshold (a and b) −log_10(p)=8.55; orange points above the line and intersected by a vertical line indicate identified SNPs with known plant height genes within ±300 kb.

Manhattan plot.

Semantic similarity for plant height spoken data GWAS

The second phenotype category used BLUEs (Fig. 6a) and BLUPs (Fig. 6b) for tall query and semantic similarity of spoken phenotype descriptions. These analyses identified 27 and 23 significant SNPs (Supplementary Table 3; colored orange), respectively, above the Bonferroni threshold of 8.55. Of these, 9 and 8 genes, respectively, were formerly detected for plant height within ±300 kb of the SNP (Table 3, Supplementary Table 2).

Fig. 6. Semantic similarity for query tall—phenotypic data. Manhattan plot generated using GAPIT and FarmCPU using semantic similarity score data a) BLUEs and b) BLUPs using Mural et al. genotypic data. The horizontal red dashed line indicates the Bonferroni threshold (a and b) −log_10(p)=8.55; orange points above the line and intersected by a vertical line indicate identified SNPs with known plant height genes within ±300 kb.

Manhattan plot.

The semantic similarity method with the tall query to generate phenotype values for each spoken observation using spaCy’s similarity function (Honnibal and Montani 2023) and the Universal Sentence Encoder (Mensio 2023) was successful. We were able to perform GWAS, and there were significant SNPs from regions associated with plant height. As this is a proof of concept study, we acknowledge that other pretrained models exist capable of calculating semantic similarity or models that can be adapted to generate similarity scores related to plant height such as BioBERT (Lee et al. 2019), those implemented by the Python gensim package (Řehůřek and Sojka 2010), or others reviewed in Koroleva et al. (2019). Further, additional queries could be employed for relating the text observations to a height value.

Binning for plant height spoken data GWAS

The third phenotype category used BLUEs (Fig. 7a) and BLUPs (Fig. 7b) for manual binning of spoken phenotype descriptions with plant height terms. These analyses identified 32 and 33, respectively, significant SNPs (Supplementary Table 3; colored orange), above the Bonferroni threshold of 8.55. Of these, 13 and 12 have genes formerly reported within ±300 kb of the SNP (Table 3, Supplementary Table 2). An additional analysis was completed using predicted values from a multinomial regression (Supplementary Fig. 8, Supplementary Table 2) in which 21 SNPs were significant, and 3 had genes detected within ±300 kb of the SNP in the literature (Table 3, Supplementary Table 2).

Fig. 7. Binned phrases for plant height—phenotypic data. Manhattan plot generated using GAPIT and FarmCPU using binned plant height data a) BLUEs and b) BLUPs using Mural et al. genotypic data. The horizontal red dashed line indicates the Bonferroni threshold (a and b) −log_10(p)=8.55; orange points above the line and intersected by a vertical line indicate identified SNPs with known plant height genes within ±300 kb.

Manhattan plot.

The binning method for plant height phrases appears to be a promising method for association studies with phenotype data extracted from spoken descriptions. This method reduces the noisiness of the transcription data and scores observations on only phrases detailing features of plant height. Additionally, the GWAS performed with BLUE and BLUP values generated from the binning method detected more known regions associated with plant height formerly reported than the manually measured and semantic similarity query methods.

Comparing across manual measurement, semantic similarity, and binning

Within the ±300 kb region indicated by the LD decay curve for each of the analyses, we lined up shared significant regions within BLUEs and BLUPs for each of the three studies (i.e. manual measurement, semantic similarity, and binning). These comparisons are shown in Supplementary Table 4. Within the manual measurements, 11 loci identified are shared across BLUEs and BLUPs. For semantic similarity, 12 shared regions are identified. For binning, 30 shared regions are identified.

Of particular interest are comparisons between manual measurement and the two methods derived from unstructured, spoken language descriptions of phenotype (i.e. semantic similarity and binning). Manual measurement and semantic similarity analyses showed no correspondence. However, manual measurement and binning shared four regions for BLUEs, with one of the four regions resulting from the same SNP (as opposed to the other three, where different SNP within the ±300 kb region are responsible for the correspondence). Four SNPs in that shared group (representing a region on Chromosome 1 and a region on Chromosome 8) contain genes known to be associated with plant height from the literature (marked in orange in Supplementary Table 4). For BLUPs, manual measurement and binning share six regions, with three regions resulting from shared SNP (and three as a result of different SNP within the ±300 kb region). Four SNPs in that shared group (representing a region on Chromosome 1 and a region on Chromosome 8) contain genes known to be associated with plant height from the literature (likewise indicated in orange). The significant shared SNPs on chromosome 8 for each group are shared across measured data binned data for both BLUEs and BLUPs.

Of some note are the regions that are not shared between manual measurements and the semantic similarity or binning GWAS outputs. For those significant associations that are only present in the semantic similarity and/or binning outputs, some of these candidate genes for plant height agree with other studies. For others, they are not documented in other studies, making them truly novel candidate genes identified via association genetics based on spoken phenotypic descriptions.

Example candidate genes in select regions

To give examples of how to review and prioritize candidate genes for the plant height trait in some of the regions identified, we describe here the three regions that were shared between the measured and binned results sets for both BLUEs and BLUPs. These constitute three loci, one each on chromosomes 3, 8, and 9. Shown in Table 4 are these three loci, of which one (on chromosome 8) coincides with a previously published association from GWAS for plant height, and the other two are newly identified through GWAS in this study from spoken phenotypic data collection techniques as well as from the manual scoring procedure.

Table 4. Candidate genes for plant height, shared measured and binned BLUEs and BLUPs.

Measured SNP ID (Position)	Binned SNP ID (position)	Gene model(s)a	Maize gene name (symbol)	GO unique IDb	GO term name	Arabidopsis ortholog(s)	
3-219332687 (222915976)	3-219281938 (222864957)	18 Total					
		Zm00001d044242	bHLH-transcription factor25 (bhlh25)	5 unrelated		AT3G21330	
		Zm00001d044255	–	6 unrelated		AT3G13882	
		Zm00001d044260	–	12 unrelated		AT5G56930	
8-665419 (801635)	c 8-872459 (1016749)	15 Total					
		Zm00001d008200	proteolipid membrane potential regulator7 (pmpm7)	GO:0009737	Response to abscisic acid	–	
				+ 4 unrelated			
							
		Zm00001d008201	aux/iaa transcription factor34 (iaa34)	GO:0009734	Auxin-activated signaling pathway	AT2G46990 AT3G17600 AT3G62100	
				+ 6 unrelated			
9-154188566 (157067037)	9-154188566 (157067037)	35 Total					
		Zm00001d048461	blue fluorescent1 (bf1)	GO:0009684	Indoleacetic acid biosynthetic process	AT5G17990	
				+ 14 unrelated			
a Only shown if any are apparently relevant to plant height trait.

b See Supplementary Table 4 for full list of Gene IDs and GO terms

cIdentified in a previously published GWAS for the trait plant height.

Table 4 shows the SNPs identified, the number of gene models in the ±300 bp region that contains the SNPs. Where characterized genes are known for those gene models, gene model names as well as gene names and symbols are also listed. GO terms relevant to the plant height trait are shown in bold. For the chromosome 3 region, no gene models in the region had any GO terms obviously related to height, so we also looked up the gene models via MaizeGDB to find out whether any associated Arabidopsis orthologs might have known phenotypes or functions related to plant height.

Promising candidate genes are as follows. For the newly identified chromosome 3 locus, three candidate gene models are: Zm00001d044242, the gene bhlh25, which has been implicated in abscisic acid biosynthesis/regulation (Vendramin et al. 2020; and Arabidopsis ortholog AT3G21330 is a member of a family of genes involved in photoresponsiveness for hypocotyl length; Khanna et al. 2006), Zm00001d044255 (Arabidopsis ortholog AT3G13882 is involved in plant growth rates and flowering; Xu et al. 2023), and Zm00001d044260, the gene c3h2, with c3h genes showing some height-related phenotypic involvement (Fornalé et al. 2015; and Arabidopsis ortholog AT5G56930, also called AtC3H65, is involved in brassinosteroid signaling; Wang et al. 2022). For the chromosome 8 locus previously associated with plant height by Azodi et al. 2020), the loci pmpm7 (also called PMP3-7 in some literature) and iaa34 are the only genes in the region. Both are known to be responsible for plant growth and height phenotypes (Fu et al. 2012; Galli et al. 2015) and are associated with GO terms that are likely relevant. For the newly identified chromosome 9 locus, one candidate gene model looks promising: Zm00001d048461 with associated term GO:0009684, indoleacetic acid biosynthetic process (where indoleacetic acid or IAA is an auxin). Upon further investigation, this gene model represents the gene blue fluorescence1 (bf1) known to affect maize plant stature (reviewed by Khavkin and Coe 1997), but not among the literature of associated plant height genes we had assembled for this work. Note: for all loci identified as being significantly associated with the trait plant height, the information that would enable these same analysis steps can be accessed via Supplementary Tables 3 and 5.

Plant growth hormone functions in genomic regions associated with plant height

To assess relationships to plant growth hormones and identified gene IDs across the whole genome, we reviewed the set of GO terms annotated to gene IDs associated with the regions ±300 kb of significant SNPs. The full list of GO terms for each model is available in Supplementary Tables 3 and 5. To examine how these terms align with plant height terms, we queried the dataset for the words auxin, brassinosteroid, and gibberellin because of their known functions in plant height regulation (Li et al. 2020 reviews the importance of these hormones). The term “auxin” was more frequently present in these datasets when compared to brassinosteroid or gibberellin (Table 5).

Table 5. Appearance of GO terms by plant hormone for each method.

Analysis	Auxin	Brassinosteroid	Gibberellin	
	Literaturea	This studyb	Literature	This study	Literature	This study	
Measured BLUEs GO Terms	2	15	0	3	0	4	
Measured BLUPs GO Terms	3	22	0	8	0	6	
Semantic Similarity Tall Query BLUEs GO Terms	12	5	0	5	0	7	
Semantic Similarity Tall Query BLUPs GO Terms	2	15	0	3	0	6	
Manual Term Bin BLUEs GO Terms	1	16	0	8	1	10	
Manual Term Bin BLUPs GO Terms	1	15	0	6	1	9	
Manual Term Bin Multinomial GO Terms	0	10	0	3	0	1	
a Appearance count for GO terms of prior literature in Table 3.

b Appearance count for GO terms unique to this study.

Additionally, other GO annotations identified in our analyses have functions that may affect plant height. Examples of these GO terms include developmental growth (GO:0048589), anatomical structure formation involved in morphogenesis (GO:0048646), and shoot system development (GO:0048367). Further, using GWAS, we found genomic regions with predicted functional annotations related to plant hormone functions that were not reported in the literature in Table 3. These regions may be of interest for follow-on experimentation to assess potential involvement in the plant height trait.

While examining GO terms, descriptions that do not relate to plant functions but were annotated to gene IDs occurred. Errant assignments of GO terms for plant-specific tasks has been described in Fattel et al. (2022). An example is the gene ID Zm00001d008201 being assigned the term animal organ development (GO:0048513). Interestingly, this ID was also assigned the terms auxin-activated signaling pathway (GO:0009734) and post-embryonic development (GO:0009791). These results demonstrate a compelling argument for reviewing GO terms carefully.

Additional methods tried and room for improvement

We demonstrate the use of a multinomial regression to generate phenotypic input with binned data, although we recognize that FarmCPU is not optimized for multinomial input. GWAS tools that utilize an ordered multinomial regression model to predict multinomial values for association studies were developed in the medical research field (German et al. 2019). Regardless, FarmCPU with multinomial binned input for plant height detected regions of the genome associated with plant height.

While participant language was not constrained, input that is less noisy and with lower data loss could be attainable if an emphasis were placed on stating specific aspects of the plant accompanied by a descriptor and reducing literary descriptive comments. An example of describing a specific aspect of a plant is “tall, green, and long” compared to “this row has tall plants”. The former statement is unclear whether the whole plant is described or a specific aspect of the plant, while the latter makes it clearer that the total plant height is described. Literary language descriptions are more difficult to compute because context is necessary to determine the meaning behind a phrase, an example, candy cane stripe. While candy cane stripe may induce a mental image of a candy cane, unless a computation model is trained to identify the literary description, the model would not be able to discern the spoken description of phenotype as a particular striped pattern.

Conclusion

By translating the linguistic complexity of human speech into data compatible with GWAS tools, we have demonstrated a successful proof of concept that paves the way for further innovation in data collection and analysis. Our 2-fold approach-semantic similarity and manual binning-proved to be robust, capturing known genetic associations with plant height and identifying new regions of interest. Of note, those describing the plants were not aware that plant height would be a focus of the study, and no guidance on what was “tall” or “short” was provided. Nonetheless, known genetic associations with the trait plant height were uncovered. These findings underscore the potential for expanding the scope of phenotypic data collection methods in genetic studies. By confirming that unstructured spoken data can yield quantifiable results for genetic association, we open the door to complementary and diverse data collection techniques that may be more tractable for nonexpert involvement. This could enable the inclusion of nonexperts in data collection, and significantly enrich the dataset by bringing in previously overlooked phenotypic nuances.

Beyond these demonstrations that spoken, unstructured phenotypic descriptions can be used to recover known associations, there are two other conceptual benefits that should be considered. Firstly, when people are describing what they see in the field rather than exclusively collecting predefined traits, the potential to uncover novel phenomena is perhaps increased. Secondly, it is the case that for many years we have used computers to analyzed structured data, so those collecting the data have limited themselves to documenting data in a structured, computer-friendly format. This is, in effect, asking people to structure their thinking and documentation like—and for—a computer. With the methods described here, the people collecting the data are enabled to think and behave in a more naturally human way for data collection. This has implications for the rate of data collection and for cognitive burden as follows. Over 3 weeks of data collection, each participant made three complete passes of the field, recording spoken observations. However, for manual scoring and measurement data, none were able to make a single complete pass of the field. Our experimental design enabled the student participants to speak and describe plant traits using their unique vocabulary and speech patterns. Participant “Zulu” reported that recording spoken observations was simpler and easier than measuring and scoring because they could make more detailed observations about different parts of the plants because recording spoken observations was both less strenuous and less mentally taxing, indicating that data collection through speech may reduce the cognitive load on field researchers.

While image-based data collection derived from unmanned vehicles equipped with sensors for visual phenotype detection continues to improve and is invaluable for certain traits (reviewed in Xiao et al. 2022), it falls short for traits requiring, e.g. tactile or olfactory observations. Coupling spoken and written language-based annotations with image analysis can enhance our understanding of complex phenotypes, harnessing human perception to capture nuanced details that purely image-based approaches might miss. This work demonstrates that such data can be collected in straightforward ways, and that phenotypic information is indeed accessible for large-scale genetics and genomics analyses.

Supplementary Material

jkae161_Supplementary_Data

Acknowledgments

We acknowledge the High-Performance Computing (HPC) facility at Iowa State University for assisting this research through computing resources and technical support. We appreciate discussions with Toni Kazic about the preliminary conceptualizations for this research project and making audio data available to assist with experimental design. We appreciate discussions with Qi Li regarding the binning methods and Kris De Brabanter for discussions about the statistics underlying GWAS. We thank Jianming Yu for helpful discussions about genomic datasets and GWAS results. We thank Brian Dilkes, Rajdeep Khangura, and Amanpreet Kaur for providing genotypic data, guidance, and examples of data processing techniques for association studies. We thank Tyler Foster and Yu-Ru Chen for enlightening discussions about GWAS methods and tools. We thank reviewers of the submitted manuscript for helpful guidance to improve the article.

Data availability

Code to recreate the analysis in this manuscript is available at CyVerse Data Commons from (Yanarella et al. 2023b) and can be accessed from: https://datacommons.cyverse.org/browse/iplant/home/shared/commons_repo/curated/Carolyn_Lawrence_Dill_Maize_WiDiv_Association_Studies_Dataset_September_2023. The deidentified spoken data described in this manuscript is exempted by Iowa State University’s Institutional Review Board (IRB ID: 21-179-00). Phenotypic data was obtained from (Yanarella et al. 2023a) and is available from: https://datacommons.cyverse.org/browse/iplant/home/shared/commons_repo/curated/Carolyn_Lawrence_Dill_Maize_WiDiv_Summer_2021_Dataset_June_2023. Genotypic data was obtained from (Mural et al. 2022b, 2022a) and can be accessed from: https://figshare.com/articles/dataset/Maize_WiDiv_SAM_1051Genotype_vcf_gz_genotype_file/19175888/1. Gene Ontology data was obtained from (Wimalanathan and Lawrence-Dill 2017) and is available from https://datacommons.cyverse.org/browse/iplant/home/shared/commons_repo/curated/Carolyn_Lawrence-Dill_maize-GAMER_maize.B73_RefGen_v4_Zm00001d.2_Oct_2017.r1.

Supplemental material available at G3 online.

Funding

This work was supported by the Iowa State University Plant Sciences Institute Faculty Scholars Award (CJLD) and the Iowa State Predictive Plant Phenomics NSF Research Traineeship (Division of Graduate Education-1545453); CJLD is a co-principal investigator, and CFY is a trainee. This research was also supported by the National Science Foundation and the United States Department of Agriculture-National Institute of Food and Agriculture AI Research Institutes program for AI Institute: for Resilient Agriculture (#2021-67021-35329) to CJLD and supporting CFY and LF.

Author contributions

CFY and CJLD conceived the idea for this project. CFY performed the analyses and drafted the initial version of the manuscript. CFY and LF evaluated the GO term analysis. All authors have read, edited, and approved the manuscript.
==== Refs
Literature cited

Abadi M , BarhamP, ChenJ, ChenZ, DavisA, DeanJ, DevinM, GhemawatS, IrvingG, IsardM, et al. 2016. Tensorflow: a system for large-scale machine learning. In: 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’16). Berkeley (CA): USENIX: The Advanced Computing Systems Association. p. 265–283.
Andorf CM , CannonEK, PortwoodJL, GardinerJM, HarperLC, SchaefferML, BraunBL, CampbellDA, VinnakotaAG, SribalusuVV, et al. 2016. Maizegdb update: new tools, data and interface for the maize model organism database. Nucleic Acids Res. 44 (D1 ):D1195–D1201. doi:10.1093/nar/gkv1007 26432828
Austin DF , LeeM. 1996. Genetic resolution and verification of quantitative trait loci for flowering and plant height with recombinant inbred lines of maize. Genome. 39 (5 ):957–968. doi:10.1139/g96-120 18469947
Azodi CB , PardoJ, VanBurenR, de los CamposG., ShiuS-H. 2020. Transcriptome-based prediction of complex traits in maize. Plant Cell. 32 (1 ):139–151. doi:10.1105/tpc.19.00332 31641024
Bai W , ZhangH, ZhangZ, TengF, WangL, TaoY, ZhengY. 2009. The evidence for non-additive effect as the main genetic component of plant height and ear height in maize using introgression line populations. Plant Breed. 129 :376–384.
Bates D , MächlerM, BolkerB, WalkerS. 2015. Fitting linear mixed-effects models using lme4. J Stat Softw. 67 (1 ):1–48. doi:10.18637/jss.v067.i01
Blakeslee JJ , PeerWA, MurphyAS. 2005. Auxin transport. Curr Opin Plant Biol. 8 (5 ):494–500. doi:10.1016/j.pbi.2005.07.014 16054428
Bradbury PJ , ZhangZ, KroonDE, CasstevensTM, RamdossY, BucklerES. 2007. TASSEL: software for association mapping of complex traits in diverse samples. Bioinformatics. 23 (19 ):2633–2635. doi:10.1093/bioinformatics/btm308 17586829
Braun IR , Lawrence-DillCJ. 2020. Automated methods enable direct computation on phenotypic descriptions for novel candidate gene prediction. Front Plant Sci. 10 :1629. doi:10.3389/fpls.2019.01629 31998331
Braun IR , YanarellaCF, Lawrence-DillCJ. 2020. Computing on phenotypic descriptions for candidate gene discovery and crop improvement. Plant Phenomics. 2020 :1963251. doi:10.34133/2020/1963251 33313544
Braun IR , YanarellaCF, RajeswariJPD, BasshamDC, Lawrence-DillCJ. 2021. The case for retaining natural language descriptions of phenotypes in plant databases and a web application as proof of concept. bioRxiv. 10.1101/2021.02.04.429796, preprint: not peer reviewed.
Brooks L , StrableJ, ZhangX, OhtsuK, ZhouR, SarkarA, HargreavesS, ElshireRJ, EudyD, PawlowskaT, et al 2009. Microdissection of shoot meristem functional domains. PLoS Genet. 5 (5 ):e1000476. doi:10.1371/journal.pgen.1000476 19424435
Carlson M . 2023. GO.db: a set of annotation maps describing the entire gene ontology. R package version 3.17.0.
Danecek P , AutonA, AbecasisG, AlbersCA, BanksE, DePristoMA, HandsakerRE, LunterG, MarthGT, SherryST, et al. 2011. The variant call format and VCFtools. Bioinformatics. 27 (15 ):2156–2158. doi:10.1093/bioinformatics/btr330 21653522
Fattel L , PsaroudakisD, YanarellaCF, ChiteriKO, DostalikHA, JoshiP, StarrDC, VuH, WimalanathanK, Lawrence-DillCJ. 2022. Standardized genome-wide function prediction enables comparative functional genomics: a new application area for Gene Ontologies in plants. GigaScience. 11 :giac023. doi:10.1093/gigascience/giac023 35426911
Fornalé S , RencoretJ, Garcia-CalvoL, CapelladesM, EncinaA, SantiagoR, RigauJ, GutiérrezA, Del RíoJ-C, Caparros-RuizD. 2015. Cell wall modifications triggered by the down-regulation of coumarate 3-hydroxylase-1 in maize. Plant Sci. 236 :272–282. doi:10.1016/j.plantsci.2015.04.007 26025540
Fu J , ZhangD-F, LiuY-H, YingS, ShiY-S, SongY-C, LiY, WangT-Y. 2012. Isolation and characterization of maize PMP3 genes involved in salt stress tolerance. PLoS One. 7 (2 ):e31101. doi:10.1371/journal.pone.0031101 22348040
Gallavotti A . 2013. The role of auxin in shaping shoot architecture. J Exp Bot. 64 (9 ):2593–2608. doi:10.1093/jxb/ert141 23709672
Galli M , LiuQ, MossBL, MalcomberS, LiW, GainesC, FedericiS, RoshkovanJ, MeeleyR, NemhauserJL, et al. 2015. Auxin signaling modules regulate maize inflorescence architecture. Proc Natl Acad Sci USA. 112 (43 ):13372–13377. doi:10.1073/pnas.1516473112 26464512
Geisler M , MurphyAS. 2005. The ABC of auxin transport: the role of p-glycoproteins in plant development. FEBS Lett. 580 (4 ):1094–1102. doi:10.1016/j.febslet.2005.11.054 16359667
German CA , SinsheimerJS, KlimentidisYC, ZhouH, ZhouJJ. 2019. Ordered multinomial regression for genetic association analysis of ordinal phenotypes at Biobank scale. Genet Epidemiol. 44 (3 ):248–260. doi:10.1002/gepi.v44.3 31879980
Goode K , ReyK. 2022. ggResidpanel: Panels and Interactive Versions of Diagnostic Plots using ’ggplot2’. R package version 0.3.0.
Hamazaki K , IwataH. 2020. RAINBOW: haplotype-based genome-wide association study using a novel SNP-set method. PLoS Comput Biol. 16 (2 ):e1007663. doi:10.1371/journal.pcbi.1007663 32059004
Hansey CN , JohnsonJM, SekhonRS, KaepplerSM, de LeonN. 2011. Genetic diversity of a maize association population with restricted phenology. Crop Sci. 51 (2 ):704–715. doi:10.2135/cropsci2010.03.0178
Hartwig T , ChuckGS, FujiokaS, KlempienA, WeizbauerR, PotluriDPV, ChoeS, JohalGS, SchulzB. 2011. Brassinosteroid control of sex determination in maize. Proc Natl Acad Sci USA. 108 (49 ):19814–19819. doi:10.1073/pnas.1108359108 22106275
Hirsch CN , FoersterJM, JohnsonJM, SekhonRS, MuttoniG, VaillancourtB, PeñagaricanoF, LindquistE, PedrazaMA, BarryK, et al. 2014. Insights into the maize pan-genome and pan-transcriptome. Plant Cell. 26 (1 ):121–135. doi:10.1105/tpc.113.119982 24488960
Honnibal M , MontaniI. 2023. spaCy v3.5.1 spancat for multi-class labeling, fixes for textcat+transformers and more. To appear.
Jansson S . 1994. The light-harvesting chlorophyll ab-binding proteins. Biochim Biophys Acta (BBA) - Bioenerg. 1184 (1 ):1–19. doi:10.1016/0005-2728(94)90148-1
Kat IP Pty Ltd (2008). WordHippo.
Kazic T . 2020. Chloe: flexible, efficient data provenance and management. bioRxiv. 10.1101/2020.01.28.923763, preprint: not peer reviewed.
Khanna R , ShenY, Toledo-OrtizG, KikisEA, JohannessonH, HwangY-S, QuailPH. 2006. Functional profiling reveals that only a small number of phytochrome-regulated early-response genes in Arabidopsis are necessary for optimal deetiolation. Plant Cell. 18 (9 ):2157–2171. doi:10.1105/tpc.106.042200 16891401
Khavkin E , CoeEH. 1997. Mapped genomic locations for developmental functions and QTLs reflect concerted groups in maize (Zea mays L.). Theor Appl Genet. 95 (3 ):343–352. doi:10.1007/s001220050569
Koroleva A , KamathS, ParoubekP. 2019. Measuring semantic similarity of clinical trial outcomes using deep pre-trained language representations. J Biomed Inform. 100 :100058. doi:10.1016/j.yjbinx.2019.100058
Lawit SJ , WychHM, XuD, KunduS, TomesDT. 2010. Maize DELLA proteins dwarf plant8 and dwarf plant9 as modulators of plant development. Plant Cell Physiol. 51 (11 ):1854–1868. doi:10.1093/pcp/pcq153 20937610
Lee J , YoonW, KimS, KimD, KimS, SoCH, KangJ. 2019. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 36 (4 ):1234–1240. doi:10.1093/bioinformatics/btz682
Lenth RV . 2023. emmeans: Estimated Marginal Means, aka Least-Squares Means. R package version 1.8.7.
Li H , WangL, LiuM, DongZ, LiQ, FeiS, XiangH, LiuB, JinW. 2020. Maize plant architecture is regulated by the ethylene biosynthetic gene ZmACS7. Plant Physiol. 183 (3 ):1184–1199. doi:10.1104/pp.19.01421 32321843
Lipka AE , TianF, WangQ, PeifferJ, LiM, BradburyPJ, GoreMA, BucklerES, ZhangZ. 2012. GAPIT: genome association and prediction integrated tool. Bioinformatics. 28 (18 ):2397–2399. doi:10.1093/bioinformatics/bts444 22796960
Liu X , HuangM, FanB, BucklerES, ZhangZ. 2016. Iterative usage of fixed and random effect models for powerful and efficient genome-wide association studies. PLoS Genet. 12 (2 ):e1005767. doi:10.1371/journal.pgen.1005767 26828793
Mazaheri M , HeckwolfM, VaillancourtB, GageJL, BurdoB, HeckwolfS, BarryK, LipzenA, RibeiroCB, KonoTJY, et al 2019. Genome-wide association analysis of stalk biomass and anatomical traits in maize. BMC Plant Biol. 19 (1 ):45. doi:10.1186/s12870-019-1653-x 30704393
Mensio M . 2023. Martinomensio/spacy-universal-sentence-encoder: Google use (universal sentence encoder) for spacy.
Merriam-Webster (2023). Merriam-Webster Online Thesaurus.
Multani DS , BriggsSP, ChamberlinMA, BlakesleeJJ, MurphyAS, JohalGS. 2003. Loss of an MDR transporter in compact stalks of maize br2 and sorghum dw3 mutants. Science. 302 (5642 ):81–84. doi:10.1126/science.1086072 14526073
Mungall CJ , GkoutosGV, SmithCL, HaendelMA, LewisSE, AshburnerM. 2010. Integrating phenotype ontologies across multiple species. Genome Biol. 11 (1 ):R2. doi:10.1186/gb-2010-11-1-r2 20064205
Mural RV , SunG, GrzybowskiM, TrossMC, JinH, SmithC, NewtonL, AndorfCM, WoodhouseMR, ThompsonAM, et al. 2022b. Association mapping across a multitude of traits collected in diverse environments in maize. GigaScience. 11 :giac080. doi:10.1093/gigascience/giac080 35997208
Mural R , SunG, GrzybowskiM, TrossMC, JinH, SmithC, NewtonL, ThompsonAM, SigmonB, SchnableJC. 2022a. Maize_WiDiv_SAM_1051Genotype.vcf.gz genotype file
Oellrich A , WallsRL, CannonEK, CannonSB, CooperL, GardinerJ, GkoutosGV, HarperL, HeM, HoehndorfR, et al. 2015. An ontology approach to comparative phenomics in plants. Plant Methods. 11 (1 ):10. doi:10.1186/s13007-015-0053-y 25774204
Peiffer JA , RomayMC, GoreMA, Flint-GarciaSA, ZhangZ, MillardMJ, GardnerCAC, McMullenMD, HollandJB, BradburyPJ, et al 2014. The genetic architecture of maize height. Genetics. 196 (4 ):1337–1356. doi:10.1534/genetics.113.159152 24514905
R Core Team (2023). R: A Language and Environment for Statistical Computing. Vienna, Austria: R Foundation for Statistical Computing.
Řehůřek R , SojkaP. 2010. Software framework for topic modelling with large corpora. In: Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks. p. 45–50, Valletta, Malta: ELRA
Reiser L , SubramaniamS, ZhangP, BerardiniT. 2022. Using the Arabidopsis information resource (TAIR) to find information about Arabidopsis genes. Curr Protoc. 2 (10 ):e574. doi:10.1002/cpz1.574 36200836
Salvi S , CornetiS, BellottiM, CarraroN, SanguinetiMC, CastellettiS, TuberosaR. 2011. Genetic dissection of maize phenology using an intraspecific introgression library. BMC Plant Biol. 11 (1 ):4. doi:10.1186/1471-2229-11-4 21211047
Sarić R , NguyenVD, BurgeT, BerkowitzO, TrtílekM, WhelanJ, LewseyMG, ČustovićE. 2022. Applications of hyperspectral imaging in plant phenotyping. Trends Plant Sci. 27 (3 ):301–315. doi:10.1016/j.tplants.2021.12.003 34998690
Stein LD . 2013. Using GBrowse 2.0 to visualize and share next-generation sequence data. Brief Bioinformatics. 14 (2 ):162–171. doi:10.1093/bib/bbt001 23376193
Sterck L . 2021. Calculate and draw custom Venn diagrams
Tang Y , LiuX, WangJ, LiM, WangQ, TianF, SuZ, PanY, LiuD, LipkaAE, et al 2016. GAPIT version 2: an enhanced integrated tool for genomic association and prediction. Plant Genome. 9 (2 ):1–9. doi:10.3835/plantgenome2015.11.0120
Teng F , ZhaiL, LiuR, BaiW, WangL, HuoD, TaoY, ZhengY, ZhangZ. 2012. ZmGA3ox2, a candidate gene for a major QTL, qPH3.1, for plant height in maize. Plant J. 73 (3 ):405–416. doi:10.1111/tpj.2013.73.issue-3 23020630
Van Rossum G , DrakeFL. 2009. Python 3 Reference Manual. Scotts Valley (CA): CreateSpace.
Venables WN , RipleyBD. 2002. Modern Applied Statistics with S. 4th ed. New York: Springer.
Vendramin S , HuangJ, CrispPA, MadzimaTF, McGinnisKM. 2020. Epigenetic regulation of ABA-induced transcriptional responses in maize. G3 Gene Genom Genet. 10 (5 ):1727–1743. doi:10.1534/g3.119.400993
Wallace JG , ZhangX, BeyeneY, SemagnK, OlsenM, PrasannaBM, BucklerES. 2016. Genome-wide association for plant height and flowering time across 15 tropical maize populations under managed drought stress and well-watered conditions in Sub-Saharan Africa. Crop Sci. 56 (5 ):2365–2378. doi:10.2135/cropsci2015.10.0632
Wang Q , SongS, LuX, WangY, ChenY, WuX, TanL, ChaiG. 2022. Hormone regulation of CCCH zinc finger proteins in plants. Int J Mol Sci. 23 (22 ):14288. doi:10.3390/ijms232214288 36430765
Wang J , ZhangZ. 2021. GAPIT version 3: boosting power and accuracy for genomic association and prediction. Genom Proteom Bioinf. 19 (4 ):629–640. doi:10.1016/j.gpb.2021.08.005
Weng J , XieC, HaoZ, WangJ, LiuC, LiM, ZhangD, BaiL, ZhangS, LiX. 2011. Genome-wide association study identifies candidate genes that affect plant height in Chinese elite maize (Zea mays L.) inbred lines. PLoS One. 6 (12 ):e29229. doi:10.1371/journal.pone.0029229 22216221
Wickham H . 2016. ggplot2: Elegant Graphics for Data Analysis. New York: Springer.
Wimalanathan K , FriedbergI, AndorfCM, Lawrence-DillCJ. 2018. Maize GO annotation—methods, evaluation, and review (maize-GAMER). Plant Direct. 2 (4 ):e00052. doi:10.1002/pld3.52 31245718
Wimalanathan K , Lawrence-DillC. 2017. maize-GAMER Annotations for maize B73 RefGen_V4 Zm00001d.2
Winkler RG , HelentjarisT. 1995. The maize dwarf3 gene encodes a cytochrome p450-mediated early step in gibberellin biosynthesis. Plant Cell. 7 (8 ):1307–1317. doi:10.1105/tpc.7.8.1307 7549486
Woodhouse MR , CannonEK, PortwoodJL, HarperLC, GardinerJM, SchaefferML, AndorfCM. 2021. A pan-genomic approach to genome databases using maize as a model system. BMC Plant Biol. 21 (1 ):385. doi:10.1186/s12870-021-03173-5 34416864
Wu A-M , RihoueyC, SevenoM, HörnbladE, SinghSK, MatsunagaT, IshiiT, LerougeP, MarchantA. 2009. The arabidopsis IRX10 and IRX10-LIKE glycosyltransferases are critical for glucuronoxylan biosynthesis during secondary cell wall formation. Plant J. 57 (4 ):718–731. doi:10.1111/tpj.2009.57.issue-4 18980649
Xiao Q , BaiX, ZhangC, HeY. 2022. Advanced high-throughput plant phenotyping techniques for genome-wide association studies: a review. J Adv Res. 35 :215–230. doi:10.1016/j.jare.2021.05.002 35003802
Xu C , SatoY, YamazakiM, BrasserM, BarbourMA, BascompteJ, ShimizuKK. 2023. Genome-wide association study of aphid abundance highlights a locus affecting plant growth and flowering in Arabidopsis thaliana. R Soc Open Sci. 10 (8 ):230399. doi:10.1098/rsos.230399 37621664
Yanarella CF , FattelL, KristmundsdóttirÁÝ, LopezMD, EdwardsJW, CampbellDA, AbelCA, Lawrence-DillCJ. 2023a. Carolyn_Lawrence_Dill_Maize_ WiDiv_Summer_2021_Dataset_June_2023
Yanarella CF , FattelL, KristmundsdóttirÁÝ, LopezMD, EdwardsJW, CampbellDA, AbelCA, Lawrence-DillCJ. 2024. Wisconsin diversity panel phenotypes: spoken descriptions of plants and supporting data. BMC Res Notes. 17 (1 ):33. doi:10.1186/s13104-024-06694-y 38263080
Yanarella CF , FattelL, Lawrence-DillCJ. 2023b. Carolyn_Lawrence_Dill_Maize_WiDiv_Association_Studies_ Dataset_September_2023
Yang W , FengH, ZhangX, ZhangJ, DoonanJH, BatchelorWD, XiongL, YanJ. 2020. Crop phenomics and high-throughput phenotyping: past decades, current challenges, and future perspectives. Mol Plant. 13 (2 ):187–214. doi:10.1016/j.molp.2020.01.008 31981735
Yao L , van de ZeddeR, KowalchukG. 2021. Recent developments and potential of robotics in plant eco-phenotyping. Emerg Topics Life Sci. 5 (2 ):289–300. doi:10.1042/ETLS20200275
Yu J , PressoirG, BriggsWH, BiIV, YamasakiM, DoebleyJF, McMullenMD, GautBS, NielsenDM, HollandJB, et al 2005. A unified mixed-model method for association mapping that accounts for multiple levels of relatedness. Nat Genet. 38 (2 ):203–208. doi:10.1038/ng1702 16380716
Zhang C , DongS-S, XuJ-Y, HeW-M, YangT-L. 2018. PopLDdecay: a fast and effective tool for linkage disequilibrium decay analysis based on variant call format files. Bioinformatics. 35 (10 ):1786–1788. doi:10.1093/bioinformatics/bty875
