Knowledge Extraction and Distillation from Large-Scale Image-Text Colonoscopy Reports Leveraging Large Language and Vision Models.
We propose to leverage text reports using large language models (LLMs) and colonoscopy images (representations) to provide pixel-level annotation of polyps thereby tackling data annotation challenges in colonoscopy.
You may also refer to the repository contributed by other co-authors.
Feb. 29th, 2024: EndoKED is under review .- If you find this work helpful, please give us a 🌟 to receive the updation.
Overview of the EndoKED design and applications to polyp diagnosis. (a) The intrinsic supervision from raw colonoscopy reports is extracted leveraging large language and vision models. The report-level lesion label is firstly extracted from the free-text description by a large language model. Then multiple instance learning (MIL) technique propagates the report-level label to the image level. The region-level bounding box is obtained from class activation map (CAM). A large vision model takes the region-level boxes as prompt and generate pixel-level lesion segmentation. (b) The image classification model for optical biopsy is developed in a data-efficient way - pre-training using multi-centre colonoscopy reports and fine-tuning with limited pathology annotation.
To clone all files:
git clone -i https://github.com/zwyang6/ENDOKED.git
To install Python dependencies:
pip install -r requirements.txt
EndoKED is evaluated on six public out-of-domain datasets, i.e., CVC-ClinicDB, Kvasir-SEG, ETIS, CVC-ColonDB, CVC-300, and the multicentre dataset PolyGen (including both sequence and frame images). EndoKED is also evaluated on the small polyp subset in PolypGen (i.e., PolypGen-Small), which is more challenging to locate and segment.
To download these six datasets, you may refer to the corresponding papers or directly download them HERE
| Dataset | Year | Resolution | Training | Testing | Total |
|---|---|---|---|---|---|
| CVC-ClinincDB | 2015 | 384x384 | 550 | 62 | 612 |
| Kvasir-SEG | 2020 | 332x487~1920x1072 | 900 | 100 | 1000 |
| ETIS | 2014 | 1225x966 | N/A | 196 | 196 |
| CVC-ColonDB | 2016 | 574x500 | N/A | 380 | 380 |
| CVC-300 | 2017 | 574x500 | N/A | 60 | 60 |
| PolypGen-Frame | 2023 | 228x384~1080x1920 | N/A | 1537 | 1537 |
| PolypGen-Video | 2023 | 576x720~1080x1920 | N/A | 1710 | 1710 |
| PolypGen-Small | 2023 | 228x384~1080x1920 | N/A | 93 | 93 |
Due to hospital confidentiality agreements, we are currently unable to release the training dataset. However, based on our EndoKED dataset, we release the checkpoint of our EndoKED-SEG model, which shows exceptional generalisation ability across the six public datasets.
<style type="text/css"> .tg {border-collapse:collapse;border-spacing:0;} .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; overflow:hidden;padding:10px 5px;word-break:normal;} .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;} .tg .tg-c3ow{border-color:inherit;text-align:center;vertical-align:top} </style>Model |
Download | Performance | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Checkpoints | Logs | Kvasir-SEG | CVC-ClinicDB | CVC-ColonDB | CVC-300 | ETIS | PolypGen-Frame | PolyGen-Video | PolyGen-Small | |
| EndoKED | full_ckpt | log | 0.883 | 0.815 | 0.806 | 0.893 | 0.819 | 0.745 | 0.668 | 0.495 |
Moreover, we have pre-trained 9 powerful baseline models (including CNN- and ViT-based architectures), i.e., Unet, Unet++, C2FNet, DCRNet, LDNet, Polyp-PVT, FCBFormer, Polyp-CASCADE, and PIDNet (Lightweigh). We have publicly released the pretrained checkpoints for further research and reproducibility.
<style type="text/css"> .tg {border-collapse:collapse;border-spacing:0;} .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; overflow:hidden;padding:10px 5px;word-break:normal;} .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;} .tg .tg-9wq8{border-color:inherit;text-align:center;vertical-align:middle} </style>| Model | Unet | Unet++ | C2FNet | DCRNet | LDNet | Polyp-PVT | FCBFormer | Polyp-CASCADE | PIDNet(lightweight) |
|---|---|---|---|---|---|---|---|---|---|
| Checkpoints | full_ckpt | full_ckpt | full_ckpt | full_ckpt | full_ckpt | full_ckpt | full_ckpt | full_ckpt | full_ckpt |
With the pre-trained segmentation models using EndoKED annotation, we combine the training set from CVC-ClinicDB and Kvasir-SEG as the final training set and evaluate its effectiveness in the testing set of the six public datasets. It demonstrate that new SOTA performance and better generalisation ability of supervised models can be achieved, with a significant gain compared with pre-training on ImageNet.
<style type="text/css"> .tg {border-collapse:collapse;border-spacing:0;} .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; overflow:hidden;padding:10px 5px;word-break:normal;} .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;} .tg .tg-c3ow{border-color:inherit;text-align:center;vertical-align:top} </style>Model |
Download | Performance | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Checkpoints | Logs | Kvasir-SEG | CVC-ClinicDB | CVC-ColonDB | CVC-300 | ETIS | PolypGen-Frame | PolyGen-Video | PolyGen-Small | |
| Unet | full_ckpt | log | 0.817 | 0.794 | 0.622 | 0.849 | 0.537 | 0.656 | 0.453 | 0.202 |
| Unet++ | full_ckpt | log | 0.833 | 0.792 | 0.616 | 0.814 | 0.524 | 0.657 | 0.454 | 0.196 |
| C2FNet | full_ckpt | log | 0.913 | 0.920 | 0.809 | 0.886 | 0.811 | 0.762 | 0.703 | 0.504 |
| DCRNet | full_ckpt | log | 0.888 | 0.901 | 0.749 | 0.870 | 0.799 | 0.742 | 0.684 | 0.593 |
| LDNet | full_ckpt | log | 0.908 | 0.905 | 0.798 | 0.902 | 0.793 | 0.747 | 0.655 | 0.420 |
| Polyp-PVT | full_ckpt | log | 0.923 | 0.937 | 0.808 | 0.900 | 0.835 | 0.777 | 0.693 | 0.585 |
| FCBFormer | full_ckpt | log | 0.914 | 0.912 | 0.812 | 0.897 | 0.821 | 0.761 | 0.646 | 0.512 |
| Polyp-CASCADE | full_ckpt | log | 0.925 | 0.936 | 0.817 | 0.905 | 0.813 | 0.769 | 0.716 | 0.561 |
| PIDNet (Lightweigh) | full_ckpt | log | 0.885 | 0.878 | 0.733 | 0.881 | 0.728 | 0.719 | 0.615 | 0.389 |
With the released checkpoints, you can utilize them for the downstream tasks.
To get started, please refer to the official codebase or choose your preferred architecture from ./EndoKED_SEG_Baselines.
Before running your baseline, download the corresponding checkpoint pretrained on the EndoKED dataset and place it under the ./logs directory.
You should also configure your dataset path in the script run_train.sh.
Finally, you can transfer EndoKED to your downstream tasks using the following command:
bash run_train.sh
1. EndoKED-MIL
pyhon ./EndoKED_MIL/train_Endo_BagDistillation_SharedEnc_Similarity_StuFilter.py
2. EndoKED-WSSS
-
bash ./EndoKED_WSSS/launch/1_data_processing.sh -
bash ./EndoKED_WSSS/launch/run_ALL.sh -
bash ./EndoKED_WSSS/launch/3_refine_CAM_2_Pseudo.sh
3. EndoKED-SEG
-
bash ./EndoKED_SEG/train.sh -
bash ./EndoKED_WSSS/launch/5_refine_Preds_2_Pseudo.sh
python ./EndoKED_WSSS/eval_tools/a1_eval_pseuo_labels_from_SAM_byPreds_fromDecoder.py
We borrowed Polyp-PVT as our segmentation model. 9 powerful baseline models, i.e., Unet, Unet++, C2FNet, DCRNet, LDNet, Polyp-PVT, FCBFormer, Polyp-CASCADE, and PIDNet (Lightweigh), are also adopted as our baselines. Segment Anything and their pre-trained weights are leveraged to refine the pseudo labels. ToCo inspires us to conduct the generation of CAMs. Many thanks to their brilliant works!
If you find this repository useful, please consider giving a star ⭐ and citation 🦖:
@article{wang2023knowledge,
title={Knowledge extraction and distillation from large-scale image-text colonoscopy records leveraging large language and vision models},
author={Wang, Shuo and Zhu, Yan and Luo, Xiaoyuan and Yang, Zhiwei and Zhang, Yizhe and Fu, Peiyao and Wang, Manning and Song, Zhijian and Li, Quanlin and Zhou, Pinghong and others},
journal={arXiv preprint arXiv:2310.11173},
year={2023}
}
If you have any question, please feel free to contact.
