Skip to content
zwyang6Public

About

[Nature BME 2025] Knowledge Extraction and Distillation from Large-Scale Image-Text Colonoscopy Reports Leveraging Large Language and Vision Models

Resources

Stars

10 stars

Watchers

2 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

Knowledge Extraction and Distillation from Large-Scale Image-Text Colonoscopy Reports Leveraging Large Language and Vision Models.

We propose to leverage text reports using large language models (LLMs) and colonoscopy images (representations) to provide pixel-level annotation of polyps thereby tackling data annotation challenges in colonoscopy.

You may also refer to the repository contributed by other co-authors.

News

  • Feb. 29th, 2024: EndoKED is under review .
  • If you find this work helpful, please give us a 🌟 to receive the updation.

Overeview

Overview of the EndoKED design and applications to polyp diagnosis. (a) The intrinsic supervision from raw colonoscopy reports is extracted leveraging large language and vision models. The report-level lesion label is firstly extracted from the free-text description by a large language model. Then multiple instance learning (MIL) technique propagates the report-level label to the image level. The region-level bounding box is obtained from class activation map (CAM). A large vision model takes the region-level boxes as prompt and generate pixel-level lesion segmentation. (b) The image classification model for optical biopsy is developed in a data-efficient way - pre-training using multi-centre colonoscopy reports and fine-tuning with limited pathology annotation.

SeCo pipeline

Dependencies

To clone all files:

git clone -i https://github.com/zwyang6/ENDOKED.git

To install Python dependencies:

pip install -r requirements.txt

Datasets

Evaluation Dataset

EndoKED is evaluated on six public out-of-domain datasets, i.e., CVC-ClinicDB, Kvasir-SEG, ETIS, CVC-ColonDB, CVC-300, and the multicentre dataset PolyGen (including both sequence and frame images). EndoKED is also evaluated on the small polyp subset in PolypGen (i.e., PolypGen-Small), which is more challenging to locate and segment.

To download these six datasets, you may refer to the corresponding papers or directly download them HERE

Dataset Year Resolution Training Testing Total
CVC-ClinincDB 2015 384x384 550 62 612
Kvasir-SEG 2020 332x487~1920x1072 900 100 1000
ETIS 2014 1225x966 N/A 196 196
CVC-ColonDB 2016 574x500 N/A 380 380
CVC-300 2017 574x500 N/A 60 60
PolypGen-Frame 2023 228x384~1080x1920 N/A 1537 1537
PolypGen-Video 2023 576x720~1080x1920 N/A 1710 1710
PolypGen-Small 2023 228x384~1080x1920 N/A 93 93

Pretrained models

EndoKED Full Checkpoints

Due to hospital confidentiality agreements, we are currently unable to release the training dataset. However, based on our EndoKED dataset, we release the checkpoint of our EndoKED-SEG model, which shows exceptional generalisation ability across the six public datasets.

<style type="text/css"> .tg {border-collapse:collapse;border-spacing:0;} .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; overflow:hidden;padding:10px 5px;word-break:normal;} .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;} .tg .tg-c3ow{border-color:inherit;text-align:center;vertical-align:top} </style>

Model
Download Performance
Checkpoints Logs Kvasir-SEG CVC-ClinicDB CVC-ColonDB CVC-300 ETIS PolypGen-Frame PolyGen-Video PolyGen-Small
EndoKED full_ckpt log 0.883 0.815 0.806 0.893 0.819 0.745 0.668 0.495

Other Baseline Models

Moreover, we have pre-trained 9 powerful baseline models (including CNN- and ViT-based architectures), i.e., Unet, Unet++, C2FNet, DCRNet, LDNet, Polyp-PVT, FCBFormer, Polyp-CASCADE, and PIDNet (Lightweigh). We have publicly released the pretrained checkpoints for further research and reproducibility.

<style type="text/css"> .tg {border-collapse:collapse;border-spacing:0;} .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; overflow:hidden;padding:10px 5px;word-break:normal;} .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;} .tg .tg-9wq8{border-color:inherit;text-align:center;vertical-align:middle} </style>
Model Unet Unet++ C2FNet DCRNet LDNet Polyp-PVT FCBFormer Polyp-CASCADE PIDNet(lightweight)
Checkpoints full_ckpt full_ckpt full_ckpt full_ckpt full_ckpt full_ckpt full_ckpt full_ckpt full_ckpt

Semantic Results

With the pre-trained segmentation models using EndoKED annotation, we combine the training set from CVC-ClinicDB and Kvasir-SEG as the final training set and evaluate its effectiveness in the testing set of the six public datasets. It demonstrate that new SOTA performance and better generalisation ability of supervised models can be achieved, with a significant gain compared with pre-training on ImageNet.

<style type="text/css"> .tg {border-collapse:collapse;border-spacing:0;} .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; overflow:hidden;padding:10px 5px;word-break:normal;} .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;} .tg .tg-c3ow{border-color:inherit;text-align:center;vertical-align:top} </style>

Model
Download Performance
Checkpoints Logs Kvasir-SEG CVC-ClinicDB CVC-ColonDB CVC-300 ETIS PolypGen-Frame PolyGen-Video PolyGen-Small
Unet full_ckpt log 0.817 0.794 0.622 0.849 0.537 0.656 0.453 0.202
Unet++ full_ckpt log 0.833 0.792 0.616 0.814 0.524 0.657 0.454 0.196
C2FNet full_ckpt log 0.913 0.920 0.809 0.886 0.811 0.762 0.703 0.504
DCRNet full_ckpt log 0.888 0.901 0.749 0.870 0.799 0.742 0.684 0.593
LDNet full_ckpt log 0.908 0.905 0.798 0.902 0.793 0.747 0.655 0.420
Polyp-PVT full_ckpt log 0.923 0.937 0.808 0.900 0.835 0.777 0.693 0.585
FCBFormer full_ckpt log 0.914 0.912 0.812 0.897 0.821 0.761 0.646 0.512
Polyp-CASCADE full_ckpt log 0.925 0.936 0.817 0.905 0.813 0.769 0.716 0.561
PIDNet (Lightweigh) full_ckpt log 0.885 0.878 0.733 0.881 0.728 0.719 0.615 0.389

Application of EndoKED-SEG

With the released checkpoints, you can utilize them for the downstream tasks. To get started, please refer to the official codebase or choose your preferred architecture from ./EndoKED_SEG_Baselines.

Before running your baseline, download the corresponding checkpoint pretrained on the EndoKED dataset and place it under the ./logs directory. You should also configure your dataset path in the script run_train.sh. Finally, you can transfer EndoKED to your downstream tasks using the following command:

bash run_train.sh

Training of EndoKED

1. EndoKED-MIL

pyhon ./EndoKED_MIL/train_Endo_BagDistillation_SharedEnc_Similarity_StuFilter.py

2. EndoKED-WSSS

  • 2.1 Data processing
    bash ./EndoKED_WSSS/launch/1_data_processing.sh
    
  • 2.2 Generating Class Activation Maps (CAMs)
    bash ./EndoKED_WSSS/launch/run_ALL.sh
    
  • 2.3 Refine CAMs to Pseudo Labels
    bash ./EndoKED_WSSS/launch/3_refine_CAM_2_Pseudo.sh
    

3. EndoKED-SEG

  • 3.1 Train EndoKED-SEG
    bash ./EndoKED_SEG/train.sh
    
  • 3.2 Refine Preds to Pseudo Labels
    bash ./EndoKED_WSSS/launch/5_refine_Preds_2_Pseudo.sh
    
  • Iterate Step 3.1-3.2 to optimize EndoKED-SEG

Evaluation of EndoKED

python ./EndoKED_WSSS/eval_tools/a1_eval_pseuo_labels_from_SAM_byPreds_fromDecoder.py

Acknowledgement

We borrowed Polyp-PVT as our segmentation model. 9 powerful baseline models, i.e., Unet, Unet++, C2FNet, DCRNet, LDNet, Polyp-PVT, FCBFormer, Polyp-CASCADE, and PIDNet (Lightweigh), are also adopted as our baselines. Segment Anything and their pre-trained weights are leveraged to refine the pseudo labels. ToCo inspires us to conduct the generation of CAMs. Many thanks to their brilliant works!

Citation

If you find this repository useful, please consider giving a star ⭐ and citation 🦖:

@article{wang2023knowledge,
  title={Knowledge extraction and distillation from large-scale image-text colonoscopy records leveraging large language and vision models},
  author={Wang, Shuo and Zhu, Yan and Luo, Xiaoyuan and Yang, Zhiwei and Zhang, Yizhe and Fu, Peiyao and Wang, Manning and Song, Zhijian and Li, Quanlin and Zhou, Pinghong and others},
  journal={arXiv preprint arXiv:2310.11173},
  year={2023}
}

If you have any question, please feel free to contact.

About

[Nature BME 2025] Knowledge Extraction and Distillation from Large-Scale Image-Text Colonoscopy Reports Leveraging Large Language and Vision Models

Resources

Stars

10 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages