Skip to content

Adapt to risk vax burden protocol - #2

Merged
marrpesce merged 14 commits into
mainfrom
adapt-to-risk-vax-burden-protocol
Oct 8, 2026
Merged

marrpesce merged 14 commits into
mainfrom
adapt-to-risk-vax-burden-protocol

Conversation

@marrpesce

Copy link
Copy Markdown
Collaborator

Hi @eparker12 and @wjchulme,

I created a first draft adapting the vaccine history scripts for the Harmonised Assessment of Risk Groups for Vaccine Prioritisation analysis.

Main changes

project.yaml

  • Added a cohort_id parameter to identify each cohort (one per target disease and time period; e.g. flu_2023_24 corresponds to the 2023/24 influenza campaign). I think this might be better than 2 parameters (e.g. target_disease + time_period), so we can run each season independently. But I am open to this approach if it is more efficient.
  • Defined three actions per cohort: extract, prepare, and report.

design.R

  • Replaced the milestone definitions with a single milestone: cohort_end_date.
  • Replaced campaign_info with cohort_info.

dataset_definition.py

  • Consolidated everything into a single dataset_definition.

  • Simplified vaccination history by removing event-level vaccination data and instead deriving two dates:

    • last vaccination before the cohort start date;
    • first vaccination after the cohort start date.
  • Removed the separate fixed dataset. We discussed with Ed that using only the latest recorded ethnicity could introduce bias, as ethnicity recording has improved over time and individuals who survive longer have more opportunities to have their ethnicity recorded. After removing ethnicity, very few truly fixed variables remained, so I incorporated them directly into the main dataset definition.

  • Added disease-specific vaccination and outcome variables, using an if statement in dataset_definition.py to generate the appropriate variables based on the cohort_id parameter.

  • I still need to add the mild outcomes. If we consider this is OK I can move to do that.

  • Age calculated the day before cohort start (cohort_start_date - days(1))

prepare.R

  • Created the baseline_vax_status variable following the protocol definitions for each vaccine.
  • Moved all variables required for the models from the report script into the preparation step.
  • Kept only vaccine product summaries to produce a simple QC table checking for any unexpected product distributions. This can easily be expanded if you think additional summaries would be useful.
  • I am not sure if it is better to collapse this script with the report_cohort.R

report_cohort.R

  • Consolidated reporting into a single script, based mainly on the former report_snapshot, and reorganised it to follow the protocol structure.
  • Added support for adjusted and unadjusted Poisson models.
  • Added a baseline vaccination history summary table.

There are still a few things to tidy up, but I think the overall structure is now much closer to the protocol and should make future extensions easier.

@wjchulme

wjchulme commented Jul 7, 2026

Copy link
Copy Markdown
Collaborator

Hi Martina, great to see this moving. Some quick thoughts from me before I dig into the code:

Added a cohort_id parameter

Nice idea.

Consolidated everything into a single dataset_definition

I think this is sensible and pragmatic. It might make the whole pipeline overall a bit slower, and total size of the outputted datasets larger, but it simplifies the code and is easier to work with. However:

only the latest recorded ethnicity could introduce bias

Unfortunately, recording dates for ethnicity are very unreliable due to patient registration movement through practices - the date of recording is lost during any transfer of records (see https://doi.org/10.1186/s12916-024-03499-5) and it defaults to 1900-01-01 (or something). So even if you choose to ignore future-dated ethnicity codes, you still might be picking up ethnicity codes that were recorded in the future (just set to 1900 by default following a move). This introduces a different type of bias (more likely to have an ethnicity code for patients who have moved in the future). Extracting ethnicity just once for each patient will also significantly decrease runtime. So something to reverting back to if speed is an issue.

Simplified vaccination history

Again, sensible and pragmatic. But it does make it difficult to identify and deal with any vaccine data quality issues of the sort we've seen in the OVERTURE work. Ideally, we would be able to create a single "cleaned_vaccinations" table view that could be queried like any other table, but we're not there yet. We can consider later whether data quality issues are concerning enough to use ELD to deal with directly.

Age calculated the day before cohort start

I would choose the day of the cohort start, because date of birth is rounded to the first of the month. If we choose the first of the month as the cohort start date, and the day before that as the age-at date, then we're essentially choosing the first of the prior month as the age-at date (i think??). Also, we chose the age-at date months after the cohort start date in ECHO, because that was how age-based vaccine eligibility worked (if you were old enough at any point during the campaign, then you were considered eligible). I don't know whether that's a sensible choice here too, but will defer to Ed.

@marrpesce
marrpesce merged commit d82e451 into main Oct 8, 2026
1 check passed
@marrpesce
marrpesce deleted the adapt-to-risk-vax-burden-protocol branch October 8, 2026 08:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants