Skip to content

Repository files navigation

HyperTex Machine Learning

This repository provides PyTorch training pipelines and inference scripts for multi-output fibre composition regression on HyperTex hyperspectral textile datasets stored in HDF5 format.


Included Tools & Files

  • train.py — Main PyTorch training script supporting:
    • Spectral MLP (--model mlp) for single-pixel spectral classification.
    • Spatial-Spectral 3D CNN (--model 3dcnn) for patch-based classification.
    • Data augmentation (flips, random Gaussian noise).
    • Validation metrics calculation (MAE, RMSE, F1-score, R², precision, recall, accuracy).
    • Model checkpoint saving (--checkpoint model.pth).
  • infer.py — Inference script for running trained model checkpoints (.pth) on test datasets or sample captures, saving prediction probability maps (.npy).
  • dataloader.py — HDF5HSIDataset PyTorch Dataset class that parses HDF5 files and produces pixel-centred samples and spatial-spectral patches.
  • requirements.txt — Python package dependencies pinned to the project environment.

Setup & Installation

Ensure you have a Python 3.9+ environment active.

If you require GPU support, install PyTorch matching your CUDA version following the official instructions first.

Then install the dependencies:

pip install -r requirements.txt

Data Sources

This repository contains software only.

Datasets are distributed separately through:

Users should download the datasets from the corresponding Zenodo records before running training or inference workflows.


HDF5 Dataset Layout

train.py, infer.py, and dataloader.py expect an HDF5 dataset file (.hdf5 or .h5) structured as follows (such as generated by hypertex-ui or supplied by HyperTex-Splits):

/
├── data/<sample_id>  # Reflectance hyperspectral cube: shape (H, W, N_BANDS)
└── gt/<sample_id>    # Fibre composition ground-truth map: shape (H, W, 10)

Model Training (train.py)

Quick Start

1. Obtain a Training Dataset

This repository operates on HDF5 datasets following the HyperTex-Splits format.

You may either:

  • Download the standardized train.hdf5 and test.hdf5 files from the HyperTex-Splits dataset.
  • Generate your own HDF5 dataset using the HyperTex UI tools from a collection of HyperTex captures and ground-truth annotations.

Store the resulting datasets in a local directory, for example:

datasets/
├── train.hdf5
└── test.hdf5

2. Spectral MLP Training (Single-Pixel)

To train a spectral MLP using single-pixel samples (patch_size = 1):

python train.py \
  --dataset /path/to/train.hdf5 \
  --model mlp \
  --patch_size 1 \
  --epochs 20 \
  --batch_size 64 \
  --lr 0.001 \
  --samples_per_image 2000 \
  --checkpoint checkpoints/model_mlp.pth \
  --cuda

3. Spatial-Spectral 3D CNN Training (Patch-Based)

To train a 3D CNN using spatial-spectral patches (patch_size > 1, e.g., 9x9):

python train.py \
  --dataset /path/to/train.hdf5 \
  --model 3dcnn \
  --patch_size 9 \
  --epochs 20 \
  --batch_size 32 \
  --lr 0.001 \
  --samples_per_image 2000 \
  --checkpoint checkpoints/model_3dcnn.pth \
  --cuda

4. Command Line Arguments for train.py

Argument Type Default Description
--dataset str Required Path to the HDF5 dataset file.
--model str mlp Architecture type: mlp or 3dcnn.
--patch_size int 1 Spatial patch size (must be 1 for mlp, >1 for 3dcnn).
--step int 1 Step size for pixel sampling window.
--samples_per_image int None Max random pixel samples to draw per image (speeds up epoch time).
--val_split float 0.2 Fraction of samples used for validation.
--batch_size int 64 Training batch size.
--epochs int 20 Number of training epochs.
--lr float 0.001 Learning rate for Adam optimizer.
--checkpoint str checkpoints/model_checkpoint.pth File path to save trained model weights.
--cuda flag False Enable CUDA GPU acceleration if available.

Inference (infer.py)

Run inference on a specific sample using a trained model checkpoint (.pth). You can supply either an HDF5 dataset file (.hdf5) or a raw capture directory:

Inference from an HDF5 dataset file:

python infer.py \
  --model_path checkpoints/model_mlp.pth \
  --model mlp \
  --dataset_path /path/to/dataset.hdf5 \
  --sample_name SAMPLE_ID \
  --patch_size 1 \
  --batch_size 64 \
  --cuda

Inference from a raw capture directory:

python infer.py \
  --model_path checkpoints/model_mlp.pth \
  --model mlp \
  --dataset_path /path/to/captures_directory \
  --sample_name SAMPLE_ID \
  --patch_size 1 \
  --batch_size 64 \
  --cuda

Features & Outputs

  • Probability output (.npy): Saves a NumPy array of shape (H, W, N_CLASSES) containing predicted pixel class probabilities (default saved to runs/<sample_name>_probs.npy or custom --output).
  • Visualization Plot (.png): Automatically generates and displays a Matplotlib figure containing an RGB preview alongside the predicted argmax classification map with color-coded fibre classes (default saved to runs/<sample_name>_inference.png or custom --plot_path). Use --no_show for headless environments.

Target Classes & Channels

Each ground-truth pixel target contains 10 continuous values representing fibre composition fractions (0.0 to 1.0):

Index Class Name
1 Unknown
2 Cotton
3 Wool
4 Lyocell
5 Viscose
6 Polyester
7 Linen
8 Elastane
9 Polyamide
10 Acrylic

Technology Stack

  • Language: Python
  • Machine Learning Framework: PyTorch
  • Data Storage: HDF5
  • Scientific Computing: NumPy
  • Visualization: Matplotlib
  • Dataset Processing: h5py

Project Status

This project is currently under active development.

Core training, evaluation, and inference workflows are stable and are used for the experiments presented in the HyperTex project. Additional models, datasets, and utilities may be introduced in future releases.


Known Issues

  • Training large spatial-spectral models may require significant GPU memory.
  • Training time depends on patch size, batch size, and sampling configuration.
  • Performance and inference speed vary according to hardware and model architecture.

License

This project is licensed under the "BSD 3-Clause License" see LICENSE for the full text.


Software and Documentation

Documentation, datasets, source code, and related resources associated with the HyperTex project are maintained through the project repositories and Zenodo records.

Links to companion datasets, source code, publications, and supplementary materials will be updated as additional resources become publicly available.


Contributing

Before contributing, please review the project governance documents:

These documents define contribution workflows, expected behaviour, and security reporting procedures.


Credits and Acknowledgements

-Tony Ferreira - Developer, INESC TEC


Contact

For questions regarding this software or the HyperTex dataset:

Tony Ferreira tony.ferreira@inesctec.pt

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages