This repository provides PyTorch training pipelines and inference scripts for multi-output fibre composition regression on HyperTex hyperspectral textile datasets stored in HDF5 format.
train.py— Main PyTorch training script supporting:- Spectral MLP (
--model mlp) for single-pixel spectral classification. - Spatial-Spectral 3D CNN (
--model 3dcnn) for patch-based classification. - Data augmentation (flips, random Gaussian noise).
- Validation metrics calculation (MAE, RMSE, F1-score, R², precision, recall, accuracy).
- Model checkpoint saving (
--checkpoint model.pth).
- Spectral MLP (
infer.py— Inference script for running trained model checkpoints (.pth) on test datasets or sample captures, saving prediction probability maps (.npy).dataloader.py—HDF5HSIDatasetPyTorch Dataset class that parses HDF5 files and produces pixel-centred samples and spatial-spectral patches.requirements.txt— Python package dependencies pinned to the project environment.
Ensure you have a Python 3.9+ environment active.
If you require GPU support, install PyTorch matching your CUDA version following the official instructions first.
Then install the dependencies:
pip install -r requirements.txtThis repository contains software only.
Datasets are distributed separately through:
- HyperTex: original hyperspectral captures, annotations, metadata, and ground-truth maps.
- Dataset: https://zenodo.org/records/21489250
- HyperTex-Splits: machine-learning-ready HDF5 training and testing partitions.
- Dataset: https://zenodo.org/records/21869714
Users should download the datasets from the corresponding Zenodo records before running training or inference workflows.
train.py, infer.py, and dataloader.py expect an HDF5 dataset file (.hdf5 or .h5) structured as follows (such as generated by hypertex-ui or supplied by HyperTex-Splits):
/
├── data/<sample_id> # Reflectance hyperspectral cube: shape (H, W, N_BANDS)
└── gt/<sample_id> # Fibre composition ground-truth map: shape (H, W, 10)
This repository operates on HDF5 datasets following the HyperTex-Splits format.
You may either:
- Download the standardized
train.hdf5andtest.hdf5files from the HyperTex-Splits dataset. - Generate your own HDF5 dataset using the HyperTex UI tools from a collection of HyperTex captures and ground-truth annotations.
Store the resulting datasets in a local directory, for example:
datasets/
├── train.hdf5
└── test.hdf5
To train a spectral MLP using single-pixel samples (patch_size = 1):
python train.py \
--dataset /path/to/train.hdf5 \
--model mlp \
--patch_size 1 \
--epochs 20 \
--batch_size 64 \
--lr 0.001 \
--samples_per_image 2000 \
--checkpoint checkpoints/model_mlp.pth \
--cudaTo train a 3D CNN using spatial-spectral patches (patch_size > 1, e.g., 9x9):
python train.py \
--dataset /path/to/train.hdf5 \
--model 3dcnn \
--patch_size 9 \
--epochs 20 \
--batch_size 32 \
--lr 0.001 \
--samples_per_image 2000 \
--checkpoint checkpoints/model_3dcnn.pth \
--cuda| Argument | Type | Default | Description |
|---|---|---|---|
--dataset |
str |
Required | Path to the HDF5 dataset file. |
--model |
str |
mlp |
Architecture type: mlp or 3dcnn. |
--patch_size |
int |
1 |
Spatial patch size (must be 1 for mlp, >1 for 3dcnn). |
--step |
int |
1 |
Step size for pixel sampling window. |
--samples_per_image |
int |
None |
Max random pixel samples to draw per image (speeds up epoch time). |
--val_split |
float |
0.2 |
Fraction of samples used for validation. |
--batch_size |
int |
64 |
Training batch size. |
--epochs |
int |
20 |
Number of training epochs. |
--lr |
float |
0.001 |
Learning rate for Adam optimizer. |
--checkpoint |
str |
checkpoints/model_checkpoint.pth |
File path to save trained model weights. |
--cuda |
flag | False |
Enable CUDA GPU acceleration if available. |
Run inference on a specific sample using a trained model checkpoint (.pth). You can supply either an HDF5 dataset file (.hdf5) or a raw capture directory:
python infer.py \
--model_path checkpoints/model_mlp.pth \
--model mlp \
--dataset_path /path/to/dataset.hdf5 \
--sample_name SAMPLE_ID \
--patch_size 1 \
--batch_size 64 \
--cudapython infer.py \
--model_path checkpoints/model_mlp.pth \
--model mlp \
--dataset_path /path/to/captures_directory \
--sample_name SAMPLE_ID \
--patch_size 1 \
--batch_size 64 \
--cuda- Probability output (
.npy): Saves a NumPy array of shape(H, W, N_CLASSES)containing predicted pixel class probabilities (default saved toruns/<sample_name>_probs.npyor custom--output). - Visualization Plot (
.png): Automatically generates and displays a Matplotlib figure containing an RGB preview alongside the predictedargmaxclassification map with color-coded fibre classes (default saved toruns/<sample_name>_inference.pngor custom--plot_path). Use--no_showfor headless environments.
Each ground-truth pixel target contains 10 continuous values representing fibre composition fractions (0.0 to 1.0):
| Index | Class Name |
|---|---|
| 1 | Unknown |
| 2 | Cotton |
| 3 | Wool |
| 4 | Lyocell |
| 5 | Viscose |
| 6 | Polyester |
| 7 | Linen |
| 8 | Elastane |
| 9 | Polyamide |
| 10 | Acrylic |
- Language: Python
- Machine Learning Framework: PyTorch
- Data Storage: HDF5
- Scientific Computing: NumPy
- Visualization: Matplotlib
- Dataset Processing: h5py
This project is currently under active development.
Core training, evaluation, and inference workflows are stable and are used for the experiments presented in the HyperTex project. Additional models, datasets, and utilities may be introduced in future releases.
- Training large spatial-spectral models may require significant GPU memory.
- Training time depends on patch size, batch size, and sampling configuration.
- Performance and inference speed vary according to hardware and model architecture.
This project is licensed under the "BSD 3-Clause License" see LICENSE for the full text.
Documentation, datasets, source code, and related resources associated with the HyperTex project are maintained through the project repositories and Zenodo records.
Links to companion datasets, source code, publications, and supplementary materials will be updated as additional resources become publicly available.
Before contributing, please review the project governance documents:
These documents define contribution workflows, expected behaviour, and security reporting procedures.
-Tony Ferreira - Developer, INESC TEC
For questions regarding this software or the HyperTex dataset:
Tony Ferreira tony.ferreira@inesctec.pt