
==== Front
Data Brief
Data Brief
Data in Brief
2352-3409
Elsevier

S2352-3409(24)00801-1
10.1016/j.dib.2024.110837
110837
Data Article
WeedCube: Proximal hyperspectral image dataset of crops and weeds for machine learning applications
Ram Billy G. a
Mettler Joseph b
Howatt Kirk b
Ostlie Michael c
Sun Xin xin.sun@ndsu.edu
a⁎
a Department of Agricultural and Biosystems Engineering, North Dakota State University, Fargo, ND, United States
b Department of Plant Science, North Dakota State University, Fargo, ND, United States
c Carrington Research Extension Center, Carrington, ND, United States
⁎ Corresponding author. xin.sun@ndsu.edu
13 8 2024
10 2024
13 8 2024
56 1108373 4 2024
9 7 2024
7 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
WeedCube dataset consists of hyperspectral images of three crops (canola, soybean, and sugarbeet) and four invasive weeds species (kochia, common waterhemp, redroot pigweed, and common ragweed). Plants were grown in two separate greenhouses and plant canopies were captured from a top-down camera angle. A push-broom hyperspectral sensor in the visible near infrared region of 400–1000 nm was used for data collection. The dataset includes 160 calibrated images. The number of images can be further increased by selection of smaller region of interests (ROIs). Dataset is supplemented by Jupyter Notebook scripts that help in data augmentation, spectral pre-processing, ROI selection for points and images, and data visualization. The primary purpose of this dataset is to support weed classification or identification studies by enhancing existing training datasets and validating the generalization capabilities of existing models. Owing to the three-dimensional (3D) nature of hyperspectral images, this dataset can also be utilized by researchers and educators across various domains for the development and testing of deep learning algorithms, the creation of automated data processing pipelines effective for 3D data, the development of tools for 3D data visualization, the creation of innovative solutions for data compression, and addressing system memory issues associated with high-dimensional data.

Keywords

Hyperspectral imaging
Deep learning
Machine learning
Weed
Crop
Precision agriculture
==== Body
pmcSpecifications TableSubject	Spectroscopy, Weed Classification, Machine Learning, Deep Learning, Hyperspectral Imaging, Proximal Sensing	
Specific subject area	Proximal hyperspectral imaging for weed classification in precision agriculture	
Type of data	Three-dimensional hyperspectral image cube of x, y, z dimension. Where, x and y are the spatial dimension and z is the spectral dimension. Provided as Numpy array (.npy) files.	
Data collection	Proximal hyperspectral image data of crops and weeds were gathered under greenhouse conditions. The data was collected at two different greenhouse locations during the years 2021–2022. The data was captured in the visible near infrared range of 400–1000 nm using a pushbroom SPECIM FX10 hyperspectral sensor mounted with a standard lens (OLET15) with 38°FOV and 150 mm focusing distance. Spatial and spectral resolution of the sensor were 1024 pixels and 5.5 nm respectively. The platform and data acquisition software used for data collection were SPECIM's LabScanner system and Lumo Scanner software, respectively. Images were taken from a top-down perspective. Halogen lights were used as the only light source. The raw hyperspectral images were calibrated using the white and dark reference images that were collected alongside each individual hyperspectral image. No data pre-processing was applied. The hyperspectral images were saved as Numpy array files (.npy), with each image stored in its own folder.	
Data source location	Data acquisition took place at two greenhouse locations.1.Waldron Greenhouse, North Dakota State University• Latitude and longitude: 46°53′42.4′′N 96°48′19.6′′W

• City/town/region: Fargo

• State: North Dakota

• Country: USA

2.Carrington Research Extension Center• Latitude and longitude: 47°30′30.0′′N 99°07′25.0′′W

• City/town/region: Carrington

• State: North Dakota

• Country: USA

	
Data accessibility	Repository name: Ag Data Commons (USDA)
Data access link: https://doi.org/10.15482/USDA.ADC/25,306,255.v1
GitHub (supplementary codes): https://github.com/billygrahamram/WeedCube	

1 Value of the Data

• The WeedCube dataset comprises hyperspectral images of three crops (canola, soybean, and sugarbeet) and four weeds (kochia, common ragweed, redroot pigweed, and common waterhemp). This data, offering both spatial and spectral information, is suitable for image or spectral based data analysis pipelines.

• Accompanying the dataset are Python scripts, saved as Jupyter notebooks, which enable viewing of pseudo RGB images from the three-dimensional (3D) data cube, plotting of spectral plots, extraction of regions of interest (ROI) as image or point spectral data, and application of data augmentation.

• The data collection for the WeedCube dataset took place in two different greenhouse located in North Dakota, USA, under controlled lighting conditions. The images captured represent various growth stages of the plants, subsequent to their vegetative phase.

• The dataset can serve as input for machine learning (ML) and deep learning (DL) models, following multivariate data analysis and statistical modelling.

• The dataset can be utilized for the development and tuning of ML and DL models that work with 3D data. It can also be used to address issues such as data compression, large data transfer, system memory exhaustion associated with high-dimensional data, preprocessing, feature selection, and the creation of data pipelines for large datasets.

• The dataset can add to the existing weed classification or identification datasets by increasing the number of training samples, serving as testing data to validate model performance, and aiding in the resolution of model overfitting for better model generalization.

• Given the multi-dimensional nature of hyperspectral images, their wider adaptability is currently limited. Our objective is to make the WeedCube dataset open source, thereby providing access to a larger community that might not have the resources to collect such data, facilitating its use for research and educational purposes.

2 Background

While traditional digital images rely primarily on spatial features of plants to distinguish crops from weeds, hyperspectral images, due to their rich spectral information, can learn from features that are less dependent on spatial arrangement [1]. This allows hyperspectral models to excel at extracting latent features that may not be readily apparent in spatial data alone [2]. In the field of precision agriculture, weed identification for site-specific weed management is a critical research area [3]. Its primary objective is to develop well-generalized models for weed detection that can be applied across diverse field conditions, including canopy structures and varying illumination patterns. WeedCube dataset [4], which provides reference-calibrated hyperspectral images of crops and weeds, can further contribute to this objective. The WeedCube dataset was compiled from two greenhouse locations during the years 2021–2022. The data was recorded in the range of 400–1000 nm. This range covers both the visible (400–700 nm) and near-infrared (700–1000 nm) wavelengths since plants interact strongly with light in these regions [5,6]. The primary objective of this dataset is to help in the creation of Machine Learning (ML) and Deep Learning (DL) models that can accurately classify crops and weeds. This objective also encompassed the creation of automated data analysis pipelines, the development of model architectures based on 3D datasets, and the validation of weed classification models. This dataset is also helpful for researchers training models on 3D dataset to address domain specific research gaps of data transfer, and memory exhaustion while ML training.

3 Data Description

This dataset provides hyperspectral images of multiple crops (soybean, canola, and sugarbeet) and weeds (kochia, common ragweed, redroot pigweed, and common waterhemp) in compressed .zip folders with descriptive filenames. Extracting the .zip files reveals individual folders containing the data (Fig. 1). Each plant folder includes hyperspectral images saved as .npy files and a corresponding pseudo RGB image saved as a .png file for visualization of data. Supplementary Jupyter Notebooks are located in the codes.zip folder and are also available in the GitHub repository for solvency of any future issues or updates. The Jupyter Notebooks allow code execution in a systematic manner and is accompanied by notes and comments for easy understanding of the code and its functions. Jupyter Notebook files include “image_roi.ipynb”, “spectral_roi.ipynb”, “augment.ipynb”, and “preprocess.ipynb”. These scripts allow further processing of the data. The images in the dataset consist of multiple plants and therefore, it would be beneficial to extract smaller ROIs for analysis purposes as 3D cube or spectral signatures. Smaller image ROIs can be selected using the “image_roi.ipynb” file which automatically saves ROIs as .npy files in the same directory. These image ROIs are 3D image cubes having x, y, z dimension. Spectral signature or point ROIs of leaves can be extracted using “spectral_roi.ipynb” and are automatically saved as .csv files. These point ROIs are 1D spectral signatures of different interest areas. Using the “augment.ipynb” file multiple data augmentation like transpose, horizontal flip, vertical flip and channel flip can be applied to the images. The “preprocess.ipynb” file can be used to pre-process the spectral signatures using Savitzky Golay first or second derivative, Standard Normal Variate (SNV), or combination of both. The “spectral_roi.ipynb” contains the mapping of the spectral dimension with actual wavelength numbers. List of images per plant can be found in Table 1.Fig. 1 The WeedCube dataset's directory structure, showcasing the organization of hyperspectral images in individual folders and Python scripts for image processing, data augmentation, and ROI selection. Each individual plant folder contains .npy hyperspectral data and .png pseudo RGB image for quick visualization.

Fig 1

Table 1 Number of samples belonging to each plant.

Table 1Plant	No. of Images	No. of plants per image	
Canola	20	4	
Soybean	20	4	
Sugarbeet	20	4	
Kochia	20	4	
Redroot pigweed	40	1	
Common Waterhemp	20	4	
Common Ragweed	20	4	

4 Experimental Design, Materials and Methods

4.1 Data acquisition

Hyperspectral data of three crops and four weeds (Fig. 2) were collected from two distinct greenhouse locations: the Waldron Greenhouse at North Dakota State University, located in Fargo, North Dakota (46°53′42.4′′N 96°48′19.6′′W) and Carrington Research Extension Center Greenhouse, located in Carrington, North Dakota (47°30′30.0′′N 99°07′25.0′′W). The data was collected in the visible near infrared region of 400–1000 nm using Specim FX10 push broom hyperspectral sensor (Spectral Imaging Ltd. Oulu, Finland). Atmospheric illumination during data collection was blocked using a light box, and the data was captured under controlled illumination from halogen light source. Crops and weeds were planted in pots (dimensions 3.5 L x 3.5 W x 5 H (inches)) in the greenhouse and were imaged while in pot. The growing medium used was off-the-self potting mix and a normal watering schedule was followed. The images captured represent various growth stages of the plants, subsequent to their vegetative phase. Each plant category includes four pots in a single image, except for redroot pigweed, which contains one plant per image.Fig. 2 Illustration of the 3D hyperspectral data in the WeedCube dataset using pseudo grayscale images created using channels 476 nm, 800 nm, and 995 nm.

Fig 2

Each hyperspectral image was captured with respective white and dark reference images for reference calibration. A white reference block of polytetrafluoroethylene (PTFE) rated to reflect 95% incident light (SphereOptics GmbH, Gewerbestraße, Germany) was imaged for white reference image and the dark reference image was captured using a closed shutter. All the images in the dataset were reference calibrated using the Eq. (1).(1) Ic=I0−IdIw−Id

Where, Ic is the calibrated spectral reflectance image, I0 is the raw spectral image, Id is the dark reference image and, Iw is the white reference image. Each plant image is a Numpy array (.npy) of dimensions x, y, z. Where x and y are the spatial dimension and z is the spectral dimension having 224 wavelengths.

4.2 Data pre-processing

The dataset is accompanied with Python scripts to generate and process data for multiple applications such as reflectance studies, spatial studies, and spatial-spectral studies.

4.3 Selection of region of interest

Region of interest refers to a user-defined subset of pixels within a hyperspectral image that represents a specific area of interest for analysis. Each pixel in a hyperspectral image contains information across a large number of wavelengths, providing a detailed spectral signature. The high dimensionality of hyperspectral data unlocks data-rich applications. While chemometric and ML approaches typically rely on one-dimensional spectral data, modern deep learning techniques, such as convolutional neural networks, can leverage the three-dimensional data cube as input [7]. Our dataset allows extraction of both spectral and spatial information using the provided Jupyter notebooks, “spectral_roi.ipynb” and “image_roi.ipynb” (Fig. 3). The ability to select multiple regions of interest (ROIs) from a single sample is particularly beneficial for data-hungry ML applications, as it effectively increases the available training data. While the dataset includes background elements (i.e., not segmented), K-means clustering has been successfully tested to achieve promising results. Users can refer Ram et al. (2023) for detailed description of using K-means clustering for background removal from hyperspectral images [8]. For user convenience, the Jupyter notebooks for spectral signature and image ROI selection have been automated. Notebooks also utilize widely popular Python libraries like Matplotlib, OpenCV, Numpy, and Pandas for easy trouble shooting. Users simply need to specify the directory containing the .npy files. Once set, the script guides users through ROI selection and saving, and automatically processes all files within the chosen folder.Fig. 3 Examples from (a) image_roi and (b) spectra_roi scripts for selection of image ROIs and point ROIs respectively.

Fig 3

4.4 Signal pre-processing of spectral data

Proximal hyperspectral data is susceptible to sensor noise and illumination variations, compromising its quality for classification problems. To address these challenges and enhance data usability, several pre-processing methods can be applied directly to the raw data (Fig. 4). Pre-processing techniques are well-established for improving the predictive power of the machine learning models [8]. The dataset includes a user-friendly Jupyter notebook, “preprocess.ipynb,” for pre-processing the raw spectral data. This script offers two key methods: the Standard Normal Variate (SNV) normalization corrects for variations in the raw data, establishing a consistent baseline for improved comparability between samples. Savitzky-Golay (SG) filtering is a smoothing technique that effectively reduces noise and enhances signal quality, particularly valuable for spectral data where peaks represent material properties. The “preprocess.ipynb” script utilizes the SciPy Signal library function for SG filtering and allows users to customize window size, polynomial order, or derivative for tailored noise reduction. Separability among classes is essential for classification problems. This can be visualized using Principal Component Analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t-SNE), and Uniform Manifold Approximation and Projection (UMAP), which are also dimensionality reduction techniques applied to hyperspectral data. PCA transforms the data to a new coordinate system, emphasizing variation and enhancing separability. t-SNE uses a probabilistic approach to model high-dimensional data by two- or three-dimensional points, preserving local structure. UMAP, similar to t-SNE, provides a balanced representation that preserves both local and global structures, facilitating better separability (Fig. 5c).Fig. 4 Pre-processing methods applied to raw spectral data of the WeedCube dataset. The x-axis represents the spectral dimension in the range of 400–1000 nm comprising of 224 wavelengths and the y-axis represents the reflectance values.

Fig 4

Fig. 5 Representation of Principal Component Analysis, t-Distributed Stochastic Neighbor Embedding, and Uniform Manifold Approximation and Projection applied to the preprocessed hyperspectral data illustrating the separability among the weed and crop spectral signatures.

Fig 5

4.5 Data augmentation

DL models require large sets of data for proper generalization capabilities. After selection of ROIs, data augmentation (Fig. 6) can be further used to increase the number of training samples available for ML or DL training [9]. Dataset consists of “augment.ipynb” that allows data augmentation of .npy files saved in the directory or smaller ROIs generated using “image_roi.ipynb”.Fig. 6 Single channel output of various augmentations that can be applied to the WeedCube dataset with the help of provided scripts.

Fig 6

4.6 Feature selection

Feature selection is the process of selecting the most significant wavelengths out of the complete spectral range of the recorded data. WeedCube data consists of 5 nm spectral resolution (Fig. 7). High spectral resolution adds redundancy to the spectral channels and increases the size of data. Feature selection methods like recursive feature elimination and principal component analysis can be used to select the most significant feature. In case of 3D image cube applications in DL models, an approach is to delete the wavelength that contain noise. Feature selections can be applied to image data and spectral data.Fig. 7 Line plot of spectral signatures of various samples from WeedCube dataset. Line plot was calculated using mean values on unprocessed point spectra.

Fig 7

Limitations

This dataset, while comprising high-quality images of various crops and weeds, has limitations to consider. DL models heavily rely on the volume of training data. While suitable for integration with larger datasets to enhance model performance, this collection contains only 160 hyperspectral images. Additionally, some images feature multiple plants, necessitating application-specific data curation (e.g., ROI selection). Furthermore, the data may exhibit occluded leaves, irregular illumination in certain instances, and sensor noise within specific spectral wavelengths.

Ethics Statement

In the context of ethical considerations, it's important to note that this paper does not involve any research conducted on human participants or animals by any of the contributing authors. The data does not originate from any online source such as search engines or social media. The dataset utilized in this study is publicly accessible and adheres to open data principles. It is incumbent upon users of these dataset to comply with standard citation practices, acknowledging the original source of the data in their work, to ensure the integrity and transparency of research processes.

CRediT authorship contribution statement

Billy G. Ram: Conceptualization, Methodology, Software, Writing – original draft. Joseph Mettler: Data curation, Writing – review & editing. Kirk Howatt: Data curation, Writing – review & editing. Michael Ostlie: . Xin Sun: Supervision, Project administration, Funding acquisition.

Data Availability

Datacube (Original data) (USDA).

Acknowledgements

This material is based upon work partially supported by the 10.13039/100000199 U.S. Department of Agriculture , agreement number 58–6064–8–023 . Any opinions, findings, conclusions, or recommendations expressed in this publication are those of the author(s) and do not necessarily reflect the view of the U.S. Department of Agriculture. This work is/was supported by the USDA 10.13039/100005825 National Institute of Food and Agriculture , Hatch project number ND01487 .

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
==== Refs
References

1 Ram B.G. A systematic review of hyperspectral imaging in precision agriculture: analysis of its current state and future prospects Comput. Electron. Agric. 222 2024 109037
2 Peña-Barragán J.M. Spectral discrimination of Ridolfia segetum and sunflower as affected by phenological stage Weed Res 46 1 2006 10 21
3 López-Granados F. Weed detection for site-specific weed management: mapping and real-time approaches Weed Res 51 1 2011 1 11
4 B.G., Ram, Proximal hyperspectral image dataset of various crops and weeds for classification via machine learning and deep learning techniques. 2024.
5 Horler D.N.H. Red edge measurements for remotely sensing plant chlorophyll content Advan. Space Res 3 2 1983 273 277
6 Bokobza L. Near infrared spectroscopy J. Near Infrared Spectrosc. 6 1 1998 3 17
7 Mishra P. Deep learning for near-infrared spectral data modelling: hypes and benefits TrAC, Trends Anal. Chem. 157 2022 116804
8 Ram B.G. Palmer amaranth identification using hyperspectral imaging and machine learning technologies in soybean field Comput. Electron. Agric. 215 2023 108444
9 Li W. Data augmentation for hyperspectral image classification with deep CNN IEEE Geosci. Remote Sens. Lett. 16 4 2019 593 597
