Skip to content

Repository files navigation

ML Pipeline

What this code does

This pipeline trains and evaluates a logistic-regression baseline and records test AUC.

Per experiment, it:

  1. Reads outcome data from outcome_folder/outcome_{outcome}.csv.
  2. Creates train/val/test split ID files in id_folder/seed{data_seed}/{outcome}/ (if missing).
  3. Reads features from feature_folder/feature.csv and aligns them to split IDs.
  4. Trains LogisticRegression in a Pipeline(StandardScaler + LogisticRegression) with GridSearchCV (roc_auc).
  5. Evaluates on test and appends summary rows to log/auc_summary.csv.

Folder structure expected

  • feature_folder/feature.csv
  • outcome_folder/outcome_{outcome}.csv
  • id_folder/seed{data_seed}/{outcome}/train_ids.csv (auto-created)
  • id_folder/seed{data_seed}/{outcome}/val_ids.csv (auto-created)
  • id_folder/seed{data_seed}/{outcome}/test_ids.csv (auto-created)

Key files

  • create_toy_data.py: generates toy feature.csv and outcome_{outcome}.csv.
  • experiments.py: experiment configuration (EXPERIMENTS, sample_id, outcome_col).
  • main.py: loops over EXPERIMENTS and writes log/auc_summary.csv.
  • run_experiment.py: one experiment run; derives outcome_file as f"outcome_{outcome}.csv".
  • save_splits.py: create train, valid, test id split.
  • read_data_split.py: reads feature/outcome/split files and builds train/val/test matrices.
  • model_pipeline.py: data standardization and model development using cross validation and hyperparameter sesarch
  • predict_test.py: use the saved best model to predict test data and evaluate the performance
  • evals_plot.py: plot auc curve and confusion matrix

Configure experiments

Edit experiments.py.

Current minimal example:

data_seed = 112
model_seed = 0
sample_id = "sample_id"
outcome_col = "event"

EXPERIMENTS = [
    {
        "outcome": "er",
        "data_seed": data_seed,
        "model_seed": model_seed,
    },
]

Notes:

  • outcome drives filename resolution: outcome_folder/outcome_{outcome}.csv.
  • sample_id and outcome_col must exist in both feature and outcome files.

Quick start (toy data)

From this directory:

python create_toy_data.py
python main.py

Outputs:

  • logs in log/
  • AUC summary in log/auc_summary.csv

Input column requirements

  • feature.csv must contain:
    • ID column: sample_id (or your configured sample_id)
    • numeric feature columns (all non-ID columns are used as model features)
  • outcome_{outcome}.csv must contain:
    • same ID column
    • outcome column named by outcome_col (default event)

About

ML pipeline for tabular classification: toy-data generation, train-test-split, sklearn Pipeline to develop models with CV and hyperparameter search

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages