This pipeline trains and evaluates a logistic-regression baseline and records test AUC.
Per experiment, it:
- Reads outcome data from outcome_folder/outcome_{outcome}.csv.
- Creates train/val/test split ID files in id_folder/seed{data_seed}/{outcome}/ (if missing).
- Reads features from feature_folder/feature.csv and aligns them to split IDs.
- Trains
LogisticRegressionin aPipeline(StandardScaler + LogisticRegression)withGridSearchCV(roc_auc). - Evaluates on test and appends summary rows to log/auc_summary.csv.
- feature_folder/feature.csv
- outcome_folder/outcome_{outcome}.csv
- id_folder/seed{data_seed}/{outcome}/train_ids.csv (auto-created)
- id_folder/seed{data_seed}/{outcome}/val_ids.csv (auto-created)
- id_folder/seed{data_seed}/{outcome}/test_ids.csv (auto-created)
create_toy_data.py: generates toy feature.csv and outcome_{outcome}.csv.experiments.py: experiment configuration (EXPERIMENTS, sample_id, outcome_col).main.py: loops over EXPERIMENTS and writes log/auc_summary.csv.run_experiment.py: one experiment run; derives outcome_file as f"outcome_{outcome}.csv".save_splits.py: create train, valid, test id split.read_data_split.py: reads feature/outcome/split files and builds train/val/test matrices.model_pipeline.py: data standardization and model development using cross validation and hyperparameter sesarchpredict_test.py: use the saved best model to predict test data and evaluate the performanceevals_plot.py: plot auc curve and confusion matrix
Edit experiments.py.
Current minimal example:
data_seed = 112
model_seed = 0
sample_id = "sample_id"
outcome_col = "event"
EXPERIMENTS = [
{
"outcome": "er",
"data_seed": data_seed,
"model_seed": model_seed,
},
]Notes:
- outcome drives filename resolution: outcome_folder/outcome_{outcome}.csv.
- sample_id and outcome_col must exist in both feature and outcome files.
From this directory:
python create_toy_data.py
python main.pyOutputs:
- logs in log/
- AUC summary in log/auc_summary.csv
- feature.csv must contain:
- ID column: sample_id (or your configured sample_id)
- numeric feature columns (all non-ID columns are used as model features)
- outcome_{outcome}.csv must contain:
- same ID column
- outcome column named by outcome_col (default event)